cover
Contact Name
Tessy Octavia Mukhti
Contact Email
tessyoctaviam@fmipa.unp.ac.id
Phone
+6282283838641
Journal Mail Official
tessyoctaviam@fmipa.unp.ac.id
Editorial Address
LPPM Universitas Negeri Padang, Jalan Prof. Dr. Hamka, Air Tawar Barat, Kota Padang, Sumatera Barat 25131
Location
Kota padang,
Sumatera barat
INDONESIA
UNP Journal of Statistics and Data Science
ISSN : -     EISSN : 2985475X     DOI : 10.24036/ujsds
UNP Journal of Statistics and Data Science is an open access journal (e-journal) launched in 2022 by Department of Statistics, Faculty of Science and Mathematics, Universitas Negeri Padang. UJSDS publishes scientific articles on various aspects related to Statistics, Data Science, and its application. Articles can be in the form of research results, case studies, or literature reviews. All papers were reviewed by peer reviewers consisting of experts and academicians across universities.
Articles 250 Documents
Classification of Stroke Desease Using the Learning Vector Quantization Algorithm Andriarmi Andriarmi; Chairina Wirdiastuti; Syafriandi Syafriandi
UNP Journal of Statistics and Data Science Vol. 4 No. 2 (2026): UNP Journal of Statistics and Data Science
Publisher : Departemen Statistika Universitas Negeri Padang

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.24036/ujsds/vol4-iss2/499

Abstract

Stroke is one of the leading causes of death and disability worldwide, thereby making early detection crucial for timely and appropriate medical treatment. In clinical practice, stroke diagnosis is generally carried out through medical examinations and patient history analysis, but this process is time-consuming and depends on the subjective judgment of medical personnel. Therefore, machine learning approaches can be utilized to support disease classification more quickly and objectively. This study aims to analyze the performance of the Learning Vector Quantization (LVQ) method in classifying stroke disease using a dataset obtained from Kaggle. The dataset used in this study is imbalanced;therefore, the SMOTE (Synthetic Minority Over-sampling Technique) method was applied to handle class imbalance. The research stages included data preprocessing, splitting data into training and testing sets, LVQ model training, parameter optimization using learning rate and maximum epoch, and model evaluation using accuracy and sensitivity. The results show that the LVQ model trained on the original dataset achieved an accuracy of 95,72%, but failed to detect stroke cases with a sensitivity of 0%. After applying SMOTE, the best model achived a stroke sensitivity of 90%, although the accuracy decreased to 49,49% due to the high number of false positives. These findings indicate that LVQ is highly sensitive to data distribution and model parameters, making its performance on this dataset less optimal for stroke classification and more suitable as an initial screening tool.
Fuzzy Time Series Singh Method for Forecasting Tourist Arrivals at Kinantan Wildlife and Cultural Park Bukittinggi Olivin Adelia Huqmi; Fadhilah Fitri; Tessy Octavia Mukhti
UNP Journal of Statistics and Data Science Vol. 4 No. 1 (2026): UNP Journal of Statistics and Data Science
Publisher : Departemen Statistika Universitas Negeri Padang

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.24036/ujsds/vol4-iss1/376

Abstract

Tourism is a key sector in regional development, contributing to economic growth, job creation, and cultural preservation. In Bukittinggi, West Sumatra, the Kinantan Wildlife and Cultural Park (TMSBK) is a major tourist destination, known for its historical and educational value. Tourist visits to TMSBK show fluctuating trends influenced by seasonal factors, socio-economic conditions, and national or global events. These dynamics make accurate forecasting essential for effective tourism planning and management. This study aims to forecast monthly tourist visits to TMSBK using the Fuzzy Time Series (FTS) Singh method, which is suitable for uncertain and fluctuating time series data. The research used historical visitor data from 2021 to 2024 obtained from the Central Bureau of Statistics. The forecasting process included defining the universe of discourse, forming class intervals, fuzzifying historical data, establishing fuzzy logical relationships (FLR), and generating forecasts. The accuracy of the forecasts was measured using Mean Absolute Percentage Error (MAPE), with a result of 19.8%, indicating good predictive performance. The results show that the FTS Singh method successfully follows the fluctuation pattern of actual visitor data. This method provides valuable insights for destination managers in planning operations, promotional efforts, and service improvements. Therefore, the FTS Singh method can be considered a reliable tool to support sustainable tourism development and decision-making in Bukittinggi.
Application of Fuzzy Time Series Cheng in Forecasting Bukittinggi's Consumer Price Index Afifah Nabilah; Fadhilah Fitri; Yenni Kurniawati
UNP Journal of Statistics and Data Science Vol. 4 No. 1 (2026): UNP Journal of Statistics and Data Science
Publisher : Departemen Statistika Universitas Negeri Padang

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.24036/ujsds/vol4-iss1/395

Abstract

The Consumer Price Index (CPI) is one of the main indicators used to measure inflation and assess the public’s purchasing power. Based on CPI monitoring in March 2025, Bukittinggi City recorded the highest year-on-year (y-o-y) inflation rate in West Sumatra at 0.50 percent, with a CPI of 106.99. This indicates significant price fluctuations, which require careful analysis and forecasting to support regional economic policymaking. This study aims to forecast the CPI of Bukittinggi City for April 2025 using the Fuzzy Time Series (FTS) Cheng method. The data used consists of monthly CPI values from January 2020 to March 2025, totaling 63 observations, obtained from the official website of Statistics Indonesia (BPS). The forecasting result using the FTS Cheng method for April 2025 shows a CPI value of 106.19. To evaluate the model's accuracy, the Mean Absolute Percentage Error (MAPE) and Mean Absolute Error (MAE) were employed, yielding values of 0.82% and 0.90%, respectively. These values fall into the “very good” category based on standard forecasting accuracy criteria. The FTS Cheng method was selected due to its ability to accommodate data fluctuations and provide weighted relationships between fuzzy intervals, thus enhancing forecasting accuracy in dynamic economic conditions. However, this study is limited to univariate data and does not compare the FTS Cheng method with other forecasting models. This research provides valuable insights for local governments in designing effective economic strategies based on reliable predictive models.
Comparison of Agglomerative Hierarchical Clustering Methods for Grouping Indonesian Provinces Based on Community Literacy Development Index Olga Afrilly Putri; Bunga Nafandra; Zamahsary Martha
UNP Journal of Statistics and Data Science Vol. 4 No. 1 (2026): UNP Journal of Statistics and Data Science
Publisher : Departemen Statistika Universitas Negeri Padang

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.24036/ujsds/vol4-iss1/470

Abstract

Community literacy development is one of the important indicators in improving the quality of human resources in Indonesia. This study aims to group provinces in Indonesia based on the Community Literacy Development Index by considering the equity of library services, the adequacy of library collections, and the level of community visits per day. The method used is agglomerative hierarchical cluster analysis. Before grouping, the data is standardized to overcome differences in units and scales between variables. The selection of the best cluster method is done using the cophenetic correlation coefficient, while the determination of the optimal number of clusters uses the silhouette method. The results of the analysis show that the Average Linkage method is the most optimal hierarchical cluster method with the best number of clusters being four clusters. Each cluster has different characteristics, reflecting variations in community literacy levels, service equity, collection adequacy, and library visit intensity between provinces. These findings indicate disparities in community literacy development between regions in Indonesia. Therefore, the results of this study are expected to serve as a basis for consideration in formulating more effective and targeted literacy and library development policies.
Spatial Autoregressive Model to Factors Poverty Gap Index in West Java, 2023 Rahmat Kurniawan; Figo Rahmatullah; Fauzan Gustiandra; Tessy Octavia Mukhti
UNP Journal of Statistics and Data Science Vol. 4 No. 1 (2026): UNP Journal of Statistics and Data Science
Publisher : Departemen Statistika Universitas Negeri Padang

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.24036/ujsds/vol4-iss1/471

Abstract

Spatial analysis is the analysis of data with spatial effects. The spatial autoregressive is used when the effect of the dependent variable at one location is influenced by the value of the dependent variable at nearby or neighboring locations. The spatial autoregressive model is more appropriate to model the factors influencing the poverty depth index in West Java in 2023. Based on the Spatial Autoregressive modeling, the variables that influence the Poverty Depth Index in West Java are Population Density, Open Unemployment Rate, and economic growth. The SAR modeling produces a higher coefficient of determination compared to the linear model, which is 68.88% with an AIC value of 18.6149.
Pengelompokan Provinsi di Indonesia Berdasarkan Indikator Pendidikan Berkualitas Tahun 2025 Menggunakan Metode Self-Organizing Maps Dinda Putri Adilla; Tessy Octavia Mukhti
UNP Journal of Statistics and Data Science Vol. 4 No. 1 (2026): UNP Journal of Statistics and Data Science
Publisher : Departemen Statistika Universitas Negeri Padang

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.24036/ujsds/vol4-iss1/472

Abstract

Access to quality education plays an essential role in improving people’s welfare and supporting sustainable development. As a fundamental component of social progress, quality education is not limited to academic attainment but also involves the development of skills, values, and character needed for meaningful participation in society. This study seeks to identify patterns and disparities in education quality among provinces by grouping regions based on multiple educational indicators. The indicators analyzed include Average Years of Schooling, Literacy Rate, Access to Information and Communication Technology, Gross Enrollment Rate, Net Enrollment Rate, and Teacher Qualifications. The data were examined using descriptive statistics, data visualization, and normalization, followed by clustering through the Self-Organizing Map (SOM) method as an unsupervised learning approach in data mining. Two clusters were formed to represent provinces with relatively higher and lower levels of educational quality. Cluster validity was assessed using internal validation measures, namely the Connectivity Index, Silhouette Index, and Dunn Index. The findings reveal that most basic education and literacy indicators show relatively favorable conditions; however, disparities remain evident in average years of schooling, ICT access, and participation in secondary and higher education. The clustering results indicate that 35 provinces fall into the group with relatively higher education quality, while 3 provinces are classified in the lower category. These results suggest that although the overall condition of education is relatively good, regional inequality in educational outcomes persists and requires targeted policy interventions to promote more balanced and inclusive development.
Evaluating Determinants Of Survival Time In Heart Failure Patients Using Cox Proportional Hazards Regression Fajri Juli Rahman Nur Zendrato; Tessy Octavia Mukhti; Sarmilah; Razita Nur Amalina; Rosa Salsabila Azarine
UNP Journal of Statistics and Data Science Vol. 4 No. 3 (2026): UNP Journal of Statistics and Data Science
Publisher : Departemen Statistika Universitas Negeri Padang

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.24036/ujsds/vol4-iss3/477

Abstract

Heart failure represents a severe chronic cardiovascular disorder and remains a major contributor to global mortality rates. Identifying specific risk factors that impact patient survivability is crucial for enhancing clinical interventions. This research investigates the determinants influencing the survival duration of heart failure patients by applying the Cox Proportional Hazards (Cox PH) regression approach. The primary objective of this study is to provide information on the main causes of heart failure and the factors that can influence the condition. Utilizing a secondary dataset of 299 patient records sourced from Kaggle, the study analyzed several variables, including age, gender, anemia, diabetes, smoking habits, and hypertension. The analytical findings reveal that age serves as the most critical determinant; individuals older than 65 face a 2.04 times greater mortality risk than younger cohorts. Furthermore, comorbidities such as anemia and hypertension significantly elevate this risk, presenting hazard ratios of 1.44 and 1.46, respectively. These outcomes emphasize the critical role of age and pre-existing medical conditions in the prognosis of heart failure, offering valuable perspectives for medical practitioners to design targeted therapeutic strategies that prolong patient survival.
Modeling Tail Risk of Employment Injury Claims using the Extreme Value Theory Framework Sri Dewi Anugrawati; Nurwahidah; Nur Aeni; Risnawati Ibnas
UNP Journal of Statistics and Data Science Vol. 4 No. 3 (2026): UNP Journal of Statistics and Data Science
Publisher : Departemen Statistika Universitas Negeri Padang

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.24036/ujsds/vol4-iss3/480

Abstract

This study investigates the implementation of Extreme Value Theory (EVT) in modeling tail risk associated with employment injury insurance claims, addressing the critical need for robust risk estimation in the presence of rare, high-severity losses. The dataset consists of 1,177 historical claims from BPJS Ketenagakerjaan Makassar recorded between 2016 and 2023. Preliminary diagnostic procedures, including hill plots and mean excess function analysis, reveal heavy-tailed behaviour with a tail index greater than one, implying the possibility of infinite variance—a characteristic often underestimated by traditional actuarial models. Due to the limited number of extreme exceedances, the block maxima framework was adopted, and a Generalized Extreme Value (GEV) distribution was fitted to weekly maxima using L-moment estimation to ensure parameter convergence. The resulting GEV parameters (location = 50.27, scale = 67.19, shape = 0.443) confirm a Fréchet-type distribution, consistent with a heavy-tailed risk profile. Comparative analysis shows that the GEV-based 99.5% VaR (1,484.09 million IDR) is more than three times higher than the empirical estimate (450.56 million IDR). Return level analysis further contextualized the estimates, indicating a one-year return level of 768.51 million IDR. These findings demonstrate that empirical methods substantially underestimate extreme loss potential, potentially threatening insurer solvency. Overall, this study provides a statistically rigorous framework for capital reserve and reinsurance optimization, offering a more conservative and theoretically justified approach to managing catastrophic occupational risks.
Monitoring The Quality of Getcontact Reviews Using Integration Sentiment Analysis and Statistical Process Control Maria Kristina Lusia Geong; Hani Brilianti Rochmanto; Fenny Fitriani; Gangga Anuraga
UNP Journal of Statistics and Data Science Vol. 4 No. 3 (2026): UNP Journal of Statistics and Data Science
Publisher : Departemen Statistika Universitas Negeri Padang

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.24036/ujsds/vol4-iss3/504

Abstract

The protection of digital communications becomes important as cybercrime rises. Getcontact, a digital communication app that offers services ranging from caller identification to spam detection, has garnered a range of opinions on Google Play Store, which serve as a basis for evaluating service quality. This study aims to monitor Getcontact’s service quality through integration of sentiment analysis and control charts based on reviews. The research data consists of reviews obtained through web scraping during the period from January 1, 2023, to September 30, 2025. The data was classified using SVM and then analyzed using p-control charts and Laney p-control charts. The results show that the negative category dominates the reviews data. SVM achieved excellent performance with accuracy 98%, precision 98%, specificity 98%, sensitivity 97%, and F1-score 97%. Control charts indicate that the process is not yet fully under control, with the Laney p chart being more representative. Pareto chart shows that complaints are dominated by issues in number identification accuracy. It is concluded that Getcontact’s service quality still needs improvement. Integration of sentiment analysis and control charts is effective for continuous quality monitoring, with the Laney p chart being more suitable.
Bayesian vs. Frequentist Estimation Under Small-Sample and Non-Normal Conditions: A Full-Factorial Monte Carlo Evaluation Aflah Zakinov Irta; Rizal Kurniawan
UNP Journal of Statistics and Data Science Vol. 4 No. 3 (2026): UNP Journal of Statistics and Data Science
Publisher : Departemen Statistika Universitas Negeri Padang

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.24036/ujsds/vol4-iss3/507

Abstract

The replication crisis in social and behavioral sciences has systematically exposed the fundamental limitations of Null Hypothesis Significance Testing (NHST). In these disciplines, small samples and non-normal data distributions are the operational norm, frequently leading to underpowered studies and fragile empirical findings. Addressing this methodological gap, we conducted a comprehensive Monte Carlo simulation (144,000 iterations) to evaluate alternative inferential frameworks under realistic research constraints. We systematically compared frequentist approaches (Welch’s t-test and Wilcoxon-Mann-Whitney) with Bayesian estimation across varied sample sizes (n = 20 to 200), effect sizes, data distributions, and prior specifications. Estimation accuracy was rigorously assessed using Mean Absolute Error (MAE), Root Mean Square Error (RMSE), and coverage rates, incorporating an explicit null condition to precisely capture Type I error dynamics. Simulation results demonstrate that under small-sample conditions typical of social and health research, Bayesian estimation utilizing an empirically derived informative prior substantially outperforms traditional frequentist models. For instance, at n = 20, the Bayesian approach reduced MAE by approximately 65% (0.124 versus 0.358) while meaningfully increasing statistical power. Critically, prior sensitivity analyses reveal that the choice of prior dictates inferential performance far more than the basic Bayesian-frequentist dichotomy. This advantage introduces a fundamental trade-off: while informative priors optimize power and minimize error in limited samples, they incrementally inflate Type I error rates. These empirically grounded findings provide behavioral researchers with actionable, condition-specific guidance for navigating inferential decisions under the realistic constraints of their discipline.