cover
Contact Name
Tessy Octavia Mukhti
Contact Email
tessyoctaviam@fmipa.unp.ac.id
Phone
+6282283838641
Journal Mail Official
tessyoctaviam@fmipa.unp.ac.id
Editorial Address
LPPM Universitas Negeri Padang, Jalan Prof. Dr. Hamka, Air Tawar Barat, Kota Padang, Sumatera Barat 25131
Location
Kota padang,
Sumatera barat
INDONESIA
UNP Journal of Statistics and Data Science
ISSN : -     EISSN : 2985475X     DOI : 10.24036/ujsds
UNP Journal of Statistics and Data Science is an open access journal (e-journal) launched in 2022 by Department of Statistics, Faculty of Science and Mathematics, Universitas Negeri Padang. UJSDS publishes scientific articles on various aspects related to Statistics, Data Science, and its application. Articles can be in the form of research results, case studies, or literature reviews. All papers were reviewed by peer reviewers consisting of experts and academicians across universities.
Articles 250 Documents
Hierarchical Bayesian Modeling with IAR Hexagonal Grids for Reconstructing Incomplete OD Matrices Eko Primadi Hendri; Sarah Fadhlia; Edi Santosa; Rachmat Sadili; Sudirman Anggada
UNP Journal of Statistics and Data Science Vol. 4 No. 3 (2026): UNP Journal of Statistics and Data Science
Publisher : Departemen Statistika Universitas Negeri Padang

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.24036/ujsds/vol4-iss3/508

Abstract

Extracting Origin-Destination (OD) matrices in open Bus Rapid Transit systems such as Transjakarta is essential for urban mobility analysis. However, this process is often hindered by incomplete observations, particularly due to missing tap-out data, which leads to extreme sparsity, zero-inflation, and overdispersion in the resulting matrices. This study addresses the problem of probabilistically reconstructing highly sparse OD matrices while accounting for spatial dependencies. To overcome limitations in previous imputation methods—such as ignoring network topology or being affected by the Modifiable Areal Unit Problem (MAUP)—this research proposes a hierarchical Bayesian approach integrating an Intrinsic Autoregressive (IAR) prior within an isotropic hexagonal (H3) tessellation framework. A Negative Binomial distribution is employed to model overdispersed count data, while latent spatial intensities and missing destinations are jointly estimated using Markov Chain Monte Carlo (MCMC). The proposed Spatial IAR model achieves stable convergence with a maximum , whereas the independent non-spatial model fails to converge adequately ( ). Although the independent model produces lower WAIC and LOOIC values (14313.41 and 14313.66) than the Spatial IAR model (14551.88 and 14611.18), the indicates that the apparent predictive superiority is spurious due to inferential instability. Posterior Predictive Checks further confirm that the spatial model successfully reproduces the overdispersion and zero-inflation characteristics of the observed mobility data. Overall, the results demonstrate that spatial regularization is essential for reconstructing high-dimensional sparse urban mobility data and improving the robustness of transportation mobility analysis.
Spatial Analysis of Earthquake Risk on Sumatra Island Using the Neyman–Scott Cox Process and Random forest Raisqa Faadillah Jatmiko; Ichsan Dwi Risky Syahputra; Fajar Rahmat; Tessy Octavia Mukhti
UNP Journal of Statistics and Data Science Vol. 4 No. 3 (2026): UNP Journal of Statistics and Data Science
Publisher : Departemen Statistika Universitas Negeri Padang

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.24036/ujsds/vol4-iss3/512

Abstract

Indonesia, located at the intersection of the Indo-Australian, Pacific, and Eurasian tectonic plates, is one of the countries most vulnerable to geological disasters, particularly earthquakes. Sumatra Island, as a densely populated region with high agricultural productivity, is highly affected by seismic activity. This study aims to analyze earthquake risk in Sumatra Island using an integrative approach that combines statistical and machine learning methods, namely the Neyman-Scott Cox Process spatial model and the Random forest algorithm. The Neyman-Scott Cox Process was used to identify spatial clustering patterns of earthquake events based on their relationship with geological features such as subduction zones, active faults, and volcanoes. In addition, Random forest was applied to develop a risk classification model based on geospatial variables and to map earthquake-prone areas into high- and medium-risk categories. The results showed that the Thomas Cluster model provided the best spatial representation with the lowest AIC value, while distance to the subduction zone had the greatest contribution to earthquake risk classification. Most western areas of Sumatra were identified as high-risk zones. The resulting zonation map can serve as a basis for decision-making in disaster mitigation policies, regional planning, and agricultural infrastructure protection. This study is expected to strengthen early warning systems, maintain food distribution and production, and contribute to food security and sustainable development in disaster-prone areas. Furthermore, this method can be replicated in other regions in Indonesia with similar geological characteristics.
Comparative Analysis of Parametric and Nonparametric Methods in Modeling Under-Five Malnutrition Prevalence Across Indonesia Dita Amelia; Suliyanto; Adelia Putri Andini; Faya Najwatus Silma; Nafla Nara Yonay; Slavina; Dinnara Chairana Aisha
UNP Journal of Statistics and Data Science Vol. 4 No. 3 (2026): UNP Journal of Statistics and Data Science
Publisher : Departemen Statistika Universitas Negeri Padang

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.24036/ujsds/vol4-iss3/554

Abstract

Malnutrition prevalence among children under five in Indonesia varied widely across provinces in 2024, ranging from 8.70% to 37.00%. Addressing this issue supports the global Sustainable Development Goals (SDGs) agenda, particularly SDG 2 (Zero Hunger) and SDG 3 (Good Health and Well-being). While most studies rely on multiple linear regression, its strict assumption of linearity is often violated in practice. This study conducts a comparative analysis between multiple linear regression and nonparametric penalized spline regression to model the effects of professional-assisted deliveries, safe sanitation access, complete basic immunization, and households at risk of stunting across 36 Indonesian provinces. Cross-sectional data for 2024 from the Central Bureau of Statistics (BPS) were analyzed, using MSE, R², and GCV for model comparison. The multiple linear regression model yielded an MSE of 18.812 and an R²  of 53.03%. Conversely, the penalized spline model (second-order polynomial, three knot points, λ = 0.0004) achieved a substantially lower MSE of 2.653 and a higher R² of 92.31%, demonstrating its superior capability in capturing nonlinear patterns. Regarding predictor significance, both models consistently identified safe sanitation, complete basic immunization, and households at risk of stunting as influential factors. However, professional assisted deliveries showed no significant effect in the nonparametric model. These findings confirm that the penalized spline approach provides more accurate estimates and is better suited for modeling provincial level malnutrition determinants.
Comparison of Least Square Spline and Penalized Spline for Modeling Human Development Index Determinants Dita Amelia; Tyo Anugrah Putra; Slavina; Thareq Alexander Manggala Napitupulu; Layyin Gisvira; Nila Khoirun Naili Salam
UNP Journal of Statistics and Data Science Vol. 4 No. 3 (2026): UNP Journal of Statistics and Data Science
Publisher : Departemen Statistika Universitas Negeri Padang

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.24036/ujsds/vol4-iss3/555

Abstract

The Human Development Index (HDI) is a key indicator of national and regional development and supports the achievement of the Sustainable Development Goals (SDGs), particularly Goal 1 (No Poverty) and Goal 4 (Quality Education). However, substantial disparities in HDI remain across Indonesia’s regencies and cities due to complex and nonlinear socioeconomic relationships that cannot always be captured by conventional parametric methods. This study analyzes the determinants of HDI in 514 regencies and cities in Indonesia in 2024 using nonparametric spline regression by comparing Penalized Spline and Least Square Spline estimators. The explanatory variables include mean years of schooling, labor force participation rate, percentage of senior-high-school graduates, and poverty rate. Secondary data from BPS were analyzed through spline basis construction, smoothing parameter selection using Generalized Cross Validation (GCV), parameter estimation, significance testing, and residual diagnostics. The results show that all predictors have nonlinear relationships with HDI. Mean years of schooling and the percentage of senior-high-school graduates positively affect HDI, whereas labor force participation and poverty rate have negative effects. The second-order Penalized Spline model with four knot points achieved the best performance, yielding the lowest GCV (5.388838), the lowest MSE (4.948574), and the highest adjusted R² (86.94%). Residual diagnostics confirmed normality and zero mean but indicated autocorrelation, suggesting spatial dependence. Overall, Penalized Spline regression provides a flexible and accurate approach for modeling HDI determinants and informing evidence-based regional development policy.
Cluster Analysis of Indonesian Provinces Based on Health Performance Indicators Using Fuzzy C-Means nurul fiskia gamayanti; Mohammad Fajri; Hartayuni Sain; Fadjryani; Iman Setiawan
UNP Journal of Statistics and Data Science Vol. 4 No. 3 (2026): UNP Journal of Statistics and Data Science
Publisher : Departemen Statistika Universitas Negeri Padang

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.24036/ujsds/vol4-iss3/529

Abstract

In accordance with the Sustainable Development Goals (SDGs), with particular emphasis on Goal 3 which focuses on ensuring a healthy life and improving the welfare of all residents across all stages of life, ranging from infancy to older adulthood, it is closely related to looking at existing health indicators. The Indonesian region consisting of 34 provinces makes policies difficult to generalize due to the different characteristics of each region. So a cluster method is needed where regional objects will be clustered into several groups so that policies will be better implemented in a formed cluster. One clustering technique that may be employed is the Fuzzy C-Means (FCM) algorithm. FCM provides a membership value in the form of degrees ranging from 0 to 1, representing the degree of strength of the relationship of a data with each group, the analysis performed using the Fuzzy C-Means (FCM) method resulted in the formation of two clusters, which represent that cluster 1 is a cluster with indications of areas that have good health, and cluster 2 is a cluster with indications of poor, Cluster 1 comprises 12 provinces, whereas Cluster 2 includes 22 provinces.
Space Time Model in Missing Value Based on Google Trends Data: Gold Price during Covid-19 Fadhlul Mubarak; Vinny Yuliani Sundara; Nurniswah; Atilla Aslanargun
UNP Journal of Statistics and Data Science Vol. 4 No. 3 (2026): UNP Journal of Statistics and Data Science
Publisher : Departemen Statistika Universitas Negeri Padang

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.24036/ujsds/vol4-iss3/534

Abstract

In certain cases, there are data that have a missing value problem. One of them is gold price data from the World Gold Council. The main purpose of this study was to predict gold prices in Austria, India, South Korea, and Turkiye during Covid-19 using the best space-time model on the data. The models that have been used in this research were generalized space-time autoregressive integrated with exogenous variable (GSTARX) and generalized space time autoregressive integrated with exogenous variable (GSTARIX). Before using these models, the last observation carried forward (LOCF) imputation technique solved the missing value problem. In addition, google trends data has been used as an alternative to the spatial weighting matrix and exogenous variables in the two models. And the google trend categories that have been used were google shopping, image search, news search, web search, and youtube search. Based on the smallest mean absolute percentage error (MAPE was 4.3%), GSTARX model in which the weighting matrix has been derived from image search. Relatively speaking, the results of forecasting gold prices in Austria was constant, India was declining, South Korea was declining significantly and Turkiye was incresing significantly.
Credit Card Fraud Detection under Extreme Class Imbalance: A Comparison of KNN and Logistic Regression Felicia Sword; Christopher Andreas
UNP Journal of Statistics and Data Science Vol. 4 No. 3 (2026): UNP Journal of Statistics and Data Science
Publisher : Departemen Statistika Universitas Negeri Padang

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.24036/ujsds/vol4-iss3/539

Abstract

Credit card fraud is a serious threat in the digital financial ecosystem and is characterised by extreme class imbalance, with fraudulent transactions typically below 1%. This study compares two standard classification algorithms, K-Nearest Neighbor (KNN) and Logistic Regression (LR), for detecting fraudulent transactions on the Sparkov dataset (1.85 million transactions; a stratified subsample of 100,000 rows; 0.52% fraud rate), and analyses the effect of the Synthetic Minority Over-sampling Technique (SMOTE). Preprocessing includes temporal feature engineering, haversine distance, leak-free per-card behavioural features, one-hot and label encoding, and z-score standardisation. Models are evaluated on a stratified 80:20 split using the confusion matrix, accuracy, precision, recall, F1-score, ROC-AUC, and PR-AUC, complemented by decision-threshold tuning, confidence intervals over five repetitions, and the McNemar test. No single model dominates across all metrics. At the default threshold, KNN baseline achieves the highest F1 (0.407) and precision (0.540), while LR baseline achieves the highest PR-AUC (0.246); LR+SMOTE leads on recall (0.712) and ROC-AUC (0.861) but with very low precision (0.025). Threshold tuning lets LR baseline reach the best F1 (0.422 at a 0.071 cut-off). McNemar shows the KNN–LR difference is not significant at baseline (p = 0.282) but significant under SMOTE (p < 0.001). The main finding is that under severe imbalance ROC-AUC can be misleading and PR-AUC is more informative; KNN baseline is a balanced detector without tuning, threshold-tuned LR baseline gives the best single operating point, and LR+SMOTE suits cases where recall is the priority.
Conditional Covariance Estimation Using CCC-GARCH for Markowitz Portfolio Optimization Indah Andirasdini; Delila Anggraini Siringo Ringo
UNP Journal of Statistics and Data Science Vol. 4 No. 3 (2026): UNP Journal of Statistics and Data Science
Publisher : Departemen Statistika Universitas Negeri Padang

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.24036/ujsds/vol4-iss3/544

Abstract

The time-varying volatility and heteroskedasticity of stock returns require appropriate volatility modeling for portfolio construction. This study aims to model stock return volatility using the Constant Conditional Correlation Generalized Autoregressive Conditional Heteroskedasticity (CCC-GARCH) and apply the estimated conditional covariance matrix to optimal portfolio construction based on the Markowitz model. The study employs daily returns of seven sector-representative stocks listed in the LQ45 Index from February 2020 to August 2025. The analysis consists of ARMA-GARCH modeling to estimate the conditional volatility of individual stock returns, CCC-GARCH estimation to obtain the conditional covariance matrix, and portfolio optimization using the Markowitz model. Portfolio performance was evaluated using the Sharpe Ratio. The results indicate that all stock returns exhibit heteroskedasticity, with the ARMA-GARCH(1,1) model providing the best volatility specification for each stock. The CCC-GARCH-based conditional covariance matrix produced an optimal portfolio with an expected return of 0.00041210, a portfolio variance of 0.00011439, and a Sharpe Ratio of 0.02103415. These findings show that the CCC-GARCH model provides a more representative estimate of portfolio risk and supports optimal portfolio construction under dynamic market conditions.
Extended Cox Model for Analyzing Factors Influencing Time to First Employment After Graduation in West Sumatra M. Anfasa Prana Karil; Zilrahmi; Rita Diana; Tessy Octavia Mukhti; Dina Fitria; Retno Lis Megawati
UNP Journal of Statistics and Data Science Vol. 4 No. 3 (2026): UNP Journal of Statistics and Data Science
Publisher : Departemen Statistika Universitas Negeri Padang

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.24036/ujsds/vol4-iss3/587

Abstract

The transition from education to employment has become one of the employment challenges in West Sumatra Province. This study aims to analyze the factors affecting the duration of obtaining a first job using the Extended Cox Proportional Hazard model. The data used were obtained from the August 2025 National Labor Force Survey (Sakernas) with variables including age, gender, educational attainment, regional classification, training, and work experience. The results show that age, educational attainment, and regional classification significantly affect the duration of obtaining a first job. Age has a positive effect that decreases over time, while higher educational attainment tends to increase job waiting time. Individuals living in rural areas tend to obtain jobs faster than those in urban areas. Meanwhile, gender and work experience are not significant, whereas training is significant through its interaction with time. Overall, the duration of obtaining a first job is influenced by individual factors, regional characteristics, and time-varying effects of the variables
Classification of Toddler Stunting Status Using Naïve Bayes Classifier with K-Fold Cross Validation Vania Riski Afifah; Zilrahmi; Syafriandi Syafriandi; Tessy Octavia Mukhti
UNP Journal of Statistics and Data Science Vol. 4 No. 3 (2026): UNP Journal of Statistics and Data Science
Publisher : Departemen Statistika Universitas Negeri Padang

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.24036/ujsds/vol4-iss3/558

Abstract

The growth of toddlers that doesn’t meet age standards can affect their quality of life in the future. In Indonesia, one of the nutritional problems that continues to receive significant attention is stunting, which is generally indicated by a mismatch between a child's height and age (height-for-age, HAZ/TB/U). Considering that determining stunting status requires a high level of accuracy, a data-driven approach is needed to support the identification and evaluation of children's nutritional conditions in a more systematic manner. This study aims to classify stunting status among toddlers using the Naïve Bayes Classifier (NBC) algorithm with a 10-Fold Cross Validation method. NBC was selected because it is simple, efficient, and suitable for probability-based classification of health data. The data were obtained from the Health Office of Pasaman Regency in 2022. The variables used include gender, birth weight, birth height, age at measurement, body weight, body height, Mid-Upper Arm Circumference (MUAC), and height-for-age (HAZ) status. Stunting status was divided into two categories: stunted and severely stunted. The analysis showed that 53.91% of toddlers were classified as stunted, while 46.10% were classified as severely stunted. Based on the evaluation using 10-Fold Cross Validation, the NBC model showed good performance with an accuracy of 86,26%, precision of 73,33%, recall of 68,75%, and an F1-score of 70,97%. The analysis also showed that height, weight, and MUAC at the time of measurement were the characteristics that most clearly distinguished the stunted and severely stunted categories. These findings indicate that anthropometric indicators can support the detection and monitoring of stunting among toddlers. Overall, this study is expected to help health workers identify toddlers who require further nutritional assessment and support more targeted stunting management efforts.