Claim Missing Document
Check
Articles

Comparison of K-Means and Ward Methods in Clustering Indonesian Provinces Based on Household Basic Service Access Nurul Mulya; Fajri Juli Rahman Nur Zendrato; Muhammad Arief Rivano; Zamahsary Martha; Tessy Octavia Mukhti
UNP Journal of Statistics and Data Science Vol. 4 No. 1 (2026): UNP Journal of Statistics and Data Science
Publisher : Departemen Statistika Universitas Negeri Padang

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.24036/ujsds/vol4-iss1/449

Abstract

Disparities in household basic service access across provinces in Indonesia remain a key issue in regional development. Basic services such as access to improved drinking water, proper sanitation, electricity, and adequate housing are essential indicators of household welfare, making regional classification necessary to identify similarities and disparities among provinces. This study aims to cluster Indonesian provinces based on household basic service access indicators and to compare the performance of the K-Means method and Hierarchical Clustering using the Ward approach. The analysis was conducted using numerical data with Euclidean distance as a measure of similarity. The optimal number of clusters was determined using the Silhouette plot and further validated using the Silhouette Coefficient. The results indicate that both K-Means and Ward methods produce two optimal clusters representing provinces with relatively high and relatively low levels of household basic service access. Centroid analysis reveals clear differences between clusters across all indicators, particularly in electricity access and sanitation. Furthermore, the evaluation of clustering quality shows that the Ward method yields a higher Silhouette Coefficient than the K-Means method, indicating more compact clusters and better separation between clusters. Therefore, the Ward method is considered more effective in mapping patterns of household basic service access across provinces. The findings of this study can support regional planning by providing a clearer understanding of disparities in household basic service access in Indonesia.
Comparison of Agglomerative Hierarchical Clustering Methods for Grouping Indonesian Provinces Based on Community Literacy Development Index Olga Afrilly Putri; Bunga Nafandra; Zamahsary Martha
UNP Journal of Statistics and Data Science Vol. 4 No. 1 (2026): UNP Journal of Statistics and Data Science
Publisher : Departemen Statistika Universitas Negeri Padang

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.24036/ujsds/vol4-iss1/470

Abstract

Community literacy development is one of the important indicators in improving the quality of human resources in Indonesia. This study aims to group provinces in Indonesia based on the Community Literacy Development Index by considering the equity of library services, the adequacy of library collections, and the level of community visits per day. The method used is agglomerative hierarchical cluster analysis. Before grouping, the data is standardized to overcome differences in units and scales between variables. The selection of the best cluster method is done using the cophenetic correlation coefficient, while the determination of the optimal number of clusters uses the silhouette method. The results of the analysis show that the Average Linkage method is the most optimal hierarchical cluster method with the best number of clusters being four clusters. Each cluster has different characteristics, reflecting variations in community literacy levels, service equity, collection adequacy, and library visit intensity between provinces. These findings indicate disparities in community literacy development between regions in Indonesia. Therefore, the results of this study are expected to serve as a basis for consideration in formulating more effective and targeted literacy and library development policies.
Implementation of Fuzzy C-Means Algorithm for Clustering Provinces in Indonesia Based on Micro and Small Industry Ratio in Village Areas Frandito Rahmanesta; Zamahsary Martha; Dodi Vionanda; Zilrahmi Zilrahmi
Indonesian Journal of Statistics and Applications Vol 8 No 2 (2024)
Publisher : Statistics and Data Science Program Study, SSMI, IPB University, in collaboration with the Forum Pendidikan Tinggi Statistika Indonesia (FORSTAT) and the Ikatan Statistisi Indonesia (ISI)

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.29244/ijsa.v8i2p178-190

Abstract

Post-economic crisis, the micro and small industries contribute the most labor compared to other industries. Regional development sourced from small micro industries is a strategic force in developing a country because the development of small micro industries leads to realizing equitable welfare to reduce income inequality. Development in village areas is an important factor for regional development, reducing inequality between regions, and alleviating poverty. However, based on the 2018 PODES survey, there are regional imbalances in Indonesia in the small micro industry which is centralized on Java Island. Therefore, clustering and characteristics of the province were carried out based on the PODES survey of the small micro industry sector. This research uses the Fuzzy C-Means algorithm to cluster 34 provinces in Indonesia based on the ratio of small micro industries in village areas in 2021, to see how the development of small micro industries in village areas in each province in Indonesia. Fuzzy C-Means is one of the data clustering techniques that uses a fuzzy clustering model, where cluster formation is based on a membership degree value that varies between 0 and 1. The Fuzzy C-Means algorithm generates 4 clusters, cluster 1 and 2 represents provinces with high and very high micro and small industry development in village areas and cluster 3 and 4 represents provinces with medium and low micro and small industry development in village areas. The Fuzzy C-Means algorithm produces a good cluster structure with a silhouette coefficient value of 0,6406.
Peramalan Curah Hujan Kabupaten Padang Pariaman dengan Menggunakan Metode Fuzzy Time Series Singh Riskiani Lubis; Zamahsary Martha; Syafriandi; Admi Salma
GAUSS: Jurnal Pendidikan Matematika Vol. 8 No. 1 (2025)
Publisher : Universitas Serang Raya

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.30656/gauss.v8i1.10465

Abstract

Abstrak Penelitian ini bertujuan untuk meramalkan curah hujan di Kabupaten Padang Pariaman, Provinsi Sumatera Barat, menggunakan metode Fuzzy Time Series Singh. Penelitian ini dilatarbelakangi oleh fluktuasi curah hujan yang tinggi di wilayah tersebut, yang menyebabkan bencana seperti banjir dan tanah longsor, yang merugikan sektor pertanian, infrastruktur, kesehatan, dan perekonomian masyarakat. Data yang digunakan adalah data curah hujan bulanan dari Januari 2020 hingga Desember 2024. Metode Fuzzy Time Series Singh dipilih karena sederhana namun efektif dalam meramalkan data runtun waktu berbasis logika fuzzy. Tahapan dalam metode ini meliputi pembentukan himpunan semesta, penentuan interval, fuzzifikasi data, pembentukan hubungan logika fuzzy, dan defuzzifikasi. Berdasarkan hasil penelitian diperoleh bahwa metode ini mampu menghasilkan estimasi curah hujan yang mendekati nilai aktual, dengan MAPE 7,67%. Hasil penelitian dapat digunakan sebagai alat bantu dalam perencanaan mitigasi bencana seperti tanah longsor dan banjir. Kata kunci: Curah Hujan, Peramalan, Fuzzy Time Series Singh Abstract This study aims to forecast rainfall in Padang Pariaman Regency, West Sumatra Province, using the Fuzzy Time Series Singh method. The research is motivated by the high fluctuation of rainfall in the area, which often leads to disasters such as floods and landslides, adversely affecting the agricultural sector, infrastructure, public health, and the local economy. The data used in this study consists of monthly rainfall records from January 2020 to December 2024. The Fuzzy Time Series Singh method was chosen due to its simplicity and effectiveness in forecasting time series data based on fuzzy logic. The stages of this method include the formation of the universe of discourse, interval determination, data fuzzification, formation of fuzzy logical relationships, and defuzzification. The results of the study show that this method is capable of producing rainfall estimates that closely match the actual values, with a MAPE of 7.67%. The findings can be used as a supporting tool for disaster mitigation planning, particularly for landslides and floods. Keywords: Rainfall, Forecasting, Fuzzy Time Series Singh
Peramalan Jumlah Curah Hujan di Kota Pariaman Menggunakan Metode ARIMA Putri, Deya Junida; Martha, Zamahsary; Salma, Admi
Journal of Authentic Research Vol. 5 No. 3 (2026): August
Publisher : LITPAM

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.36312/jar.v5i3.6451

Abstract

Curah hujan merupakan salah satu unsur iklim yang berpengaruh terhadap berbagai sektor, seperti pertanian, perikanan, kesehatan, transportasi serta pengelolaan sumber daya di Indonesia. Kota Pariaman yang merupakan wilayah pesisir dengan tingkat kerentanan banjir yang memerlukan sistem peramalan jumlah curah hujan yang akurat sebagai antisipasi dampak bencana yang ditimbulkan dan untuk mendukung perencanaan pembangunan. Penelitian ini bertujuan untuk melakukan peramalan data deret waktu jumlah curah hujan di Kota Pariaman dengan menggunakan metode Autoregressive Integrated Moving Average (ARIMA). Penelitian ini menggunakan data sekunder yang diambil dari website Badan Pusat Statistik (BPS) Kota Pariaman terkait jumlah curah hujan dari Januari 2020 sampai Desember 2024. Berdasarkan hasil analisis, diperoleh model terbaik yaitu  model ARIMA (2,0,2) dengan nilai MSE terkecil dan residual yang berdistribusi normal. Hasil analisis juga menunjukkan nilai MAPE sebesar 6,81% artinya model ARIMA (2,0,2) sudah sangat akurat digunakan. Hasil peramalan ini diharapkan dapat dijadikan dasar pengambilan keputusan bagi pemerintah dan masyarakat dalam mengantisipasi dampak bencana dan mendukung perencanaan pembangunan di Kota Pariaman. Rainfall is one of the climate elements that affect various sectors, such as agriculture, fisheries, health, transportation and resource management in Indonesia. Pariaman City, which is a coastal area that has a high vulnerability to flood disasters, requires an accurate rainfall forecasting system to anticipate the impact of disasters caused and to support development planning. This study aims to forecast time series data on the amount of rainfall in Pariaman city using the Autoregressive Integrated Moving Average (ARIMA) method. This research uses secondary data taken from the Pariaman City Statistics Agency (BPS) website regarding the amount of rainfall from January 2020 to December 2024. Based on the results of the analysis, the best model is the ARIMA (2,0,2) model with the smallest MSE value and normally distributed residuals. The analysis results also show a MAPE value of 6.81%, meaning that the ARIMA (2,0,2) model is very accurate to use. The results of this forecasting are expected to be used as a basis for decision making for the supporting development planning in Pariaman City.
Penerapan Spatial Autoregressive Model pada Kasus Multikolinearitas (Studi Kasus: Indeks Ketahanan Pangan di Indonesia) Diah Uswatun Hasanah; Zamahsary Martha
TSAQOFAH Vol 6 No 5 (2026): TSAQOFAH: Jurnal Penelitian Guru Indonesia
Publisher : Lembaga Yasin AlSys

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.58578/tsaqofah.v6i5.11813

Abstract

Food security modeling in Indonesia needs to consider interregional linkages and strong correlations among predictor variables. However, the application of spatial models that simultaneously address spatial dependence and multicollinearity at the provincial level remains limited. This study aimed to apply the Spatial Autoregressive Model (SAR) to model Indonesia’s 2024 Food Security Index (FSI), using Principal Component Analysis (PCA) as an approach to address multicollinearity. The study employed an applied quantitative approach using secondary data from 38 provinces obtained from Badan Pangan Nasional and Badan Pusat Statistik. Seven predictor variables were analyzed using multiple linear regression, variance inflation factor (VIF), PCA, Moran’s I, Lagrange Multiplier, and SAR with a k-nearest neighbors (KNN) spatial weight matrix. The results indicated high multicollinearity among three predictor variables. PCA produced three principal components that explained 94.50% of the variance in the data. A Moran’s I value of 0.71005767 indicated positive spatial autocorrelation, while the SAR spatial coefficient of 0.52676 indicated a significant spatial effect. The SAR model performed better than multiple linear regression, with an AIC value of 242.5907 and an R² of 0.7577. These findings confirm that integrating PCA and SAR produces FSI modeling that is more consistent with the characteristics of the data and can support the formulation of food security policies that consider interprovincial linkages. Future research may develop a spatial panel approach to analyze FSI dynamics over time.