Claim Missing Document
Check
Articles

CLUSTERING MODEL K-MEANS PADA KASUS ANGKA PUTUS SEKOLAH TINGKATAN SEKOLAH DASAR DI PROVINSI JAWA TENGAH Laila Khoirun Nisa; Tari Fitri Ningsih; Burhanuddin Izzul Salam; Fauzi, Fatkhurokhman; Eny Winaryati
LogicLink Vol. 1 No. 1, June 2024
Publisher : Universitas Islam Negeri K.H. Abdurrahman Wahid Pekalongan

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.28918/logiclink.v1i1.7793

Abstract

Basic education aims to equip children with the basic skills they need to navigate their lives as individuals, elements of society, elements of citizens and also as human beings, and prepare them for higher education in the future. The case of dropping out of school seems to be a problem that cannot be overcome. The impact of dropping out of school if not managed properly will certainly be detrimental, including having an impact on the quality of resources in the future. Therefore, it is necessary to take action to reduce the dropout rate at the elementary school level. This research was conducted using the K-Means clustering algorithm to find out which districts/cities in Central Java have high, medium, and low dropout rates. The results of clustering using the K-Means algorithm through 3 methods obtained an optimal K value of 3, therefore 3 clusters were formed from all 35 regencies/cities in Central Java, cluster 1 with a tendency for high elementary school dropout rates to be 14 regencies/cities, cluster 2 with a moderate trend of dropout rates from elementary school there are 15 regencies/cities, and cluster 3 with a low trend of dropout rates from elementary school there are 6 regencies/cities.
OPTIMALKAN PEMAHAMAN DATA DENGAN DASHBOARD MELALUI PELATIHAN VISUALISASI DATA UNTUK SISWA SMA N 1 KEMBANG JEPARA Fadlurohman, Alwan; Fauzi, Fatkhurokhman; Lestari, Febi Anggun; Sarah, Albertus Dion
Community Development Journal : Jurnal Pengabdian Masyarakat Vol. 5 No. 5 (2024): Vol. 5 No. 5 Tahun 2024
Publisher : Universitas Pahlawan Tuanku Tambusai

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.31004/cdj.v5i5.34893

Abstract

Seiring dengan kemajuan teknologi informasi dan memasuki era revolusi industri 4.0, pemanfaatan teknologi dalam aktivitas manusia semakin meningkat. Data kini menjadi aset berharga, dan kemampuan dalam mengumpulkan, menganalisis, serta menginterpretasi data telah menjadi keterampilan krusial dalam dunia kerja dan pendidikan. Pendidikan memainkan peran penting dalam mempersiapkan generasi mendatang untuk menghadapi tantangan global yang semakin kompleks. Di SMAN 1 Kembang Kabupaten Jepara, terdapat beberapa masalah utama, yaitu keterbatasan akses dan pemahaman siswa mengenai data, kurangnya pengalaman dalam visualisasi data, dan minimnya pemahaman tentang manfaat visualisasi data. Program pelatihan ini bertujuan untuk meningkatkan pemahaman siswa mengenai data dan teknik visualisasi melalui penyampaian materi konseptual tentang konsep data dan visualisasi, penggunaan Google Data Studio, dan pembuatan dashboard visualisasi. Pelaksanaan PKM ini menggunakan metode interaktif dan demontrasi kepada siswa kelas 11 SMAN 1 Kembang sebanyak 20 orang. Hasil PKM menunjukkan bahwa siswa mampu meningkatkan pemahaman mereka tentang konsep dasar data, jenis-jenis data, dan pentingnya visualisasi data. Siswa juga menunjukkan kemajuan yang baik dalam keterampilan teknis terkait penggunaan Google Data Studio untuk mengolah dan memvisualisasikan data. Mereka berhasil menerapkan berbagai fitur untuk menyusun grafik, diagram, dan visualisasi lain yang sesuai dengan data yang diberikan. Hasil visualisasi ini memperlihatkan kemampuan siswa dalam menyajikan data dengan cara yang informatif dan menarik.
Klasterisasi Indikator Kesehatan Ibu dan Anak di Indonesia Menggunakan Hierarchical Clustering Agglomerative Angelina, Lea; Putri, Dinda Meyda; Ana, Nisfatun Nurul; Syafira, Elsa Izza; Chumairoh, Kamilah Citra; Syaharani, Nabbila Dyah; Fauzi, Fatkhurokhman
Seminar Nasional Official Statistics Vol 2025 No 1 (2025): Seminar Nasional Official Statistics 2025
Publisher : Politeknik Statistika STIS

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.34123/semnasoffstat.v2025i1.2401

Abstract

Maternal and child health is a top priority in national development, given that high maternal and infant mortality rates remain a significant challenge in Indonesia. Disparities in health indicators between regions indicate that existing inequalities remain insufficiently addressed. This study aims to group 38 provinces in Indonesia based on 14 maternal and child health indicators for 2024 to identify patterns of disparity. The method used is Agglomerative Hierarchical Clustering with multicollinearity tests (VIF) and KMO for data validation. The complete linkage method was selected for its optimal performance, yielding an agglomerative coefficient of 0.742 and the highest silhouette value of 0.2099 at K = 6. The results formed six clusters reflecting similarities in regional characteristics. Several provinces in Papua clustered separately due to their low health indicator achievements. These findings emphasize the need for region-specific intervention policies to address disparities and promote equitable improvements in maternal and child health.
Comparison Analysis of Hierarchical Clustering and K-Means Methods in Grouping Provinces in Indonesia Based on Dengue Hemorrhagic Fever (DHF) Cases Alfidha Rahmah; Nida Faoziatun Khusna; Safril Ahmadi Sanmas; Syifa Aulia; Shinta Amaria; Fatkhurokhman Fauzi
JUITA: Jurnal Informatika JUITA Vol. 13 Issue 2, July 2025
Publisher : Department of Informatics Engineering, Universitas Muhammadiyah Purwokerto

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.30595/juita.v13i2.26131

Abstract

Indonesia, a tropical country, experiences climate variations that influence the spread of infectious diseases, including Dengue Hemorrhagic Fever (DHF). The increase in DHF cases necessitates clustering provinces based on their vulnerability to design effective mitigation strategies. This study compares two clustering methods: Hierarchical Clustering and K-Means Clustering. Within the hierarchical clustering analysis, five linkage methods were evaluated: Average Linkage, Complete Linkage, Single Linkage, Ward’s Method, and Centroid Linkage. The best linkage method was identified using the cophenetic correlation coefficient, indicating that Average Linkage produced the most representative cluster structure, resulting in three distinct groups. For the K-Means method, the optimal number of clusters was determined using the Silhouette Coefficient, which also indicated three clusters. Clustering performance evaluation revealed that Average Linkage outperformed K-Means, with a higher Silhouette Score of 0.552. The resulting clusters categorized provinces into three risk groups: high-risk areas (e.g., DKI Jakarta), moderate-risk areas (e.g., West Java and East Java), and low-risk areas, comprising the remaining provinces in Indonesia
HYBRID RESAMPLING METHOD AND HYPERPARAMETER OPTIMIZATION FOR HIV/AIDS PREDICTION: EVIDENCE FROM EIGHT MACHINE-LEARNING MODELS Lydia Nur Sa'adah; Fatkhurokhman Fauzi; Prizka Rismawati Arum; M Al Haris; Yan Nazala Bisoumi
JITK (Jurnal Ilmu Pengetahuan dan Teknologi Komputer) Vol. 11 No. 4 (2026): JITK Issue May 2026
Publisher : LPPM Nusa Mandiri

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.33480/jitk.v11i4.7533

Abstract

HIV/AIDS remains a global health challenge with continuously increasing infection rates, highlighting the importance of accurate prediction models to support prevention and early detection. However, the development of such models is often constrained by class imbalance and irrelevant features. This study aims to improve HIV/AIDS infection prediction by integrating feature selection, data balancing techniques, and eight machine learning algorithms. Feature selection was performed using Mutual Information and Chi-Square to identify the most relevant features. The dataset used was the HIV/AIDS Infection Prediction Dataset from Kaggle, consisting of 2,139 instances and 23 features, with an imbalanced distribution of 1,618 non-infected and 521 infected cases. The dataset was divided into 80% training data and 20% testing data, with resampling applied only to the training set to prevent data leakage. Three resampling scenarios were evaluated: no sampling, SMOTE, and SMOTE-ENN. Hyperparameter tuning was conducted using Bayesian Optimization integrated with 5-fold Cross-Validation to improve model robustness and reliability. Eight machine learning algorithms were evaluated, including Decision Tree, Random Forest, AdaBoost, Gradient Boosting, XGBoost, LightGBM, K-Nearest Neighbors, and Logistic Regression. The results show that SMOTE-ENN combined with hyperparameter optimization significantly improved model performance. The best model, Gradient Boosting + SMOTE-ENN, achieved 96.1% accuracy, 94.8% precision, 98.4% recall, and 96.5% F1-score. These findings indicate that the proposed integrated framework is highly effective for predicting HIV/AIDS infection and has strong potential to support early diagnosis and data-driven decision-making in healthcare.
Perbandingan Hasil Klasifikasi Data Iris menggunakan Algoritma K-Nearest Neighbor dan Random Forest : Comparison of Iris Data Classification Results using the K-Nearest Neighbor and Random Forest Algorithms Budiono Rahman; Fatkhurokhman Fauzi; Saeful Amri
Journal of Data Insights Vol 1 No 1 (2023): Journal of Data Insights
Publisher : Department of Sains Data UNIMUS Universitas Muhammadiyah Semarang

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.26714/jodi.v1i1.135

Abstract

Data mining merupakan suatau metode yang baik untuk menangani data skala besar. Performasi menjadi penting dalam metode data mining. Dua metode yang memiliki performasi terbaik diantaranya K-Nearest Neighbor (KNN) dan Random Forest (RF). Artikel ini membahas terkait perbandingan performasi K-NN dan RF. Data yang digunakan pada penelitian ini adalah Iris. Data dibagi menjadi 80% data training dan 20% data testing. Validasi performasi menggunakan nilai akurasi dan F1-Score. Berdasarkan nilai. Berdasarkan hasil yang didapat metode RF lebih baik dibandingkan dengan metode K-NN. Nilai akurasi yang didapat oleh metode RF adalah 1.00 atau 100% dan nilai F1-Score sebesar 1.00.
Clustering Untuk Menentukan Indeks Kesejahteraan Rakyat di Provinsi Jawa Tengah 2022 Menggunakan Metode Fuzzy C-Means: Clustering to Determine the People's Welfare Index in Central Java Province 2022 Using the Fuzzy C-Means Method Indra Firmansyah; Salmaa Fauziah; Hanif Nur Ibrahim; Fatkhurokhman Fauzi; Tiani Wahyu Utami
Journal of Data Insights Vol 1 No 2 (2023): Journal of Data Insights
Publisher : Department of Sains Data UNIMUS Universitas Muhammadiyah Semarang

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.26714/jodi.v1i2.149

Abstract

Kesejahtraan rakyat merupakan salah satu tujuan negara yang tercantum pada Undang-undang Dasar 1945. Dalam meningkatakan kesejahtraan rakyat, tentunya perlu adanya pembangunan yang merata. Untuk menjalankan program pembangunan yang merata, harus dilakukan identifikasi berdsarkan karaktaeristik tingkat kesejahtraan rakyat berdasarkan variabel-variabel yang telah ditentukan agar dalam membuat strategi dan mengambil kebijakan untuk meningkatkan kesejahtraan rakyat dapat tepat sasaran dan optimal. Tujuan dari penelitian ini adalah untuk mengetahui pengelompokkan 35 Kabupatan/Kota di Provinsi Jawa Tengah dan k arakteristik dari setiap kelompok berdasarkan indeks kesejahtraan rakyat. Berdasarkan hasil analisis yang dilakukan, bahwa terdapat 35 Kabupaten/Kota di Provinsi Jawa Tengah dapat membentuk 4 kelompok (cluster), dimana pada cluster 0 beranggotakan 8 Kabupaten/Kota dengan karakteristik Jumlah peduduk miskin tinggi, Daya beli cenderung rendah, rata-rata lama sekolah rendah, angka harapan hidup sangat rendah. Pada cluster 1 terdapat 12 Kabupaten/Kota dengan karakteristik Nilai PDRB sangat tinggi, angka pengangguran relatif rendah, angka lama sekolah relatif tinggi. Pada cluster 2 terdapat 5 Kabupaten/Kota dengan karakteristik Nilai PDRB sangat rendah, jumlah penduduk miskin sangat rendah, daya beli sangat tinggi, kepemilikan rumah rendah, kepadatan penduduk tinggi, daya beli tinggi, angka melek huruf tinggi, rata-rata lama sekolah tinggi, dan yang terakhir cluster 3 terdapat 10 Kabupaten/Kota dengan karakteristik PDRB rendah, kepadatan penduduk sangat rendah, kepemilikan rumah sangat tinggi, daya beli cenderung rendah.
Decision Tree Classification Prediction of Covid-19 Cases in Indonesia: Prediksi Kasus Covid-19 di Indonesia Menggunakan Metode Klasifikasi Decision Tree Amaliah Sholeha Arafat; Aprilla Anawai Basman; Fatkhurokhman Fauzi; Saeful Amri
Journal of Data Insights Vol 2 No 2 (2024): Journal of Data Insights
Publisher : Department of Sains Data UNIMUS Universitas Muhammadiyah Semarang

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.26714/jodi.v2i2.213

Abstract

Forecasting is the prediction of an event in the present and future using past event data. The purpose of forecasting is to minimize errors in predictions (forecast errors) to provide a higher level of confidence. In the context of the COVID-19 pandemic, forecasting the number of cases can help anticipate surges, allowing for better-preparedness to minimize its impact. Forecasting methods can be categorized into three common classifications: qualitative methods, time series, and causal methods. Time series methods are further divided into statistical methods and machine learning. Machine learning methods are more effective in forecasting as they can accommodate non-linear and complex relationships between inputs and outputs. One of the machine learning methods used is the Decision Tree, which is a predictive model structured in a tree or hierarchical format. The Decision tree is a data processing method for predicting the future by constructing classification and regression models in a tree structure. The decision tree is also the most popular and easily understood classification method. In this study, a classification decision tree is used to forecast positive COVID-19 cases in Indonesia using the Python programming language.
K-Nearest Neighbor (KNN) Method for Weather Data Prediction: Penerapan Metode K-Nearest Neighbour (KNN) Untuk Prediksi Data Cuaca Agata Dwi Putri Putri; M. Al Haris; Fatkhurokhman Fauzi; Saeful Amri
Journal of Data Insights Vol 3 No 1 (2025): Journal of Data Insights
Publisher : Department of Sains Data UNIMUS Universitas Muhammadiyah Semarang

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.26714/jodi.v3i1.214

Abstract

The weather tends to change frequently every day, so weather forecasts are made to be used as an early warning if sudden weather changes occur. By forecasting the weather, losses can be minimized and people are alert to carry out outdoor activities. From this problem, the K-Nearest Neighbor (KNN) method was applied. This method is expected to provide accurate and efficient information to obtain weather predictions for existing conditions. The data used is secondary data. After conducting research on training data (old data) amounting to 80% and test data (new data) amounting to 20%. The accuracy results from the testing data predictions are 75% with a value of k = 8.
Analysis Autocorrelation Spatial on Amount Fundraising at LAZISMU Semarang City Using Moran's Index: Analisis Autokorelasi Spasial pada Jumlah Penghimpunan Dana di LAZISMU Kota Semarang Menggunakan Indeks Moran Choirunnisa Hasna Nisa; Khansa' Ni'mal 'Abidah; M. Al Haris; Fatkhurokhman Fauzi
Journal of Data Insights Vol 3 No 2 (2025): Journal of Data Insights
Publisher : Department of Sains Data UNIMUS Universitas Muhammadiyah Semarang

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.26714/jodi.v3i2.314

Abstract

Institution Zakat and Infaq Collectors And Sed e kah Muhammadiyah (LAZISMU) , has role important in gather And distribute funds activity social use help communities in need . L AZISMU Semarang City in general special focus on management funds at the level city , with not quite enough answer gather And allocate funds from public to humanitarian programs like help education , health , and help social research​ This aim For increase effectiveness collection funds Institution Zakat, Infaq , and Charity Collectors Alms Muhammadiyah in Semarang City. With apply approach spatial , research This analyze pattern distribution geographical donors , potential donations , and characteristics economy as well as demographics in each sub-district . Methodology study involving spatial data collection and analysis statistics . Results study This expected can give contribution on understanding scientific related zakat- based management spatial And become guidelines for institution similar in optimize collection And allocation funds .
Co-Authors - Tarsan Achmad Fauzan Achmad Fauzan, Achmad Agata Dwi Putri Putri Agung Subakti Nuzulullail Ahmad Amrullah B Alfidha Rahmah Alfidha Rahmah Alifia Puspita Sari Alwan Fadlurohman Amaliah Sholeha Arafat Amri, Saeful Amrullah, Setiawan Ana, Nisfatun Nurul Angelina, Lea Anggun Erya Santika Anisa Ma'u Luthfi Anita Retno Indriani Aprilla Anawai Basman Ardelita Ika Fadhlillah Ardiansyah, Muhammad Rifqy Aulia, Syifa Azizah, Apipah Nur Bagus Sartono Barlian, Seftia Amelia Rizki Budiono Rahman Burhanuddin Izzul Salam Choirunnisa Hasna Nisa Chumairoh, Kamilah Citra Dannu Purwanto Dewi Ratnasari Wijaya Dwi Agustina Eko Yuliyanto, Eko Eny Winaryati Erika Siva Aulia Erlinda, Relly Fabiola, Gwenda Fadillah, Muhammad Reza Fatikha Adha Fahreza Fauziah Rahma Gabriella Hilary Wenur Hanif Nur Ibrahim Haris, M Al Haris, M. Al Hasbi Assidiqi Iis Widya Harmoko Iis Widya Harmoko Iis Widya Harmoko, Iis Widya Indah Manfaati Nur Indah Manfaati Nur Indra Firmansyah Iqbal Kharisudin Izzah, Nasyiatul Junaidi, Muhammad Rifki Khansa' Ni'mal 'Abidah Khikman, Muhammad Alvaro Laila Khoirun Nisa Lestari, Febi Anggun Lia Aryanti Sholekhah Litasya Shofwatillah Lydia Nur Sa'adah M. Al Haris Moh Yamin Darsyah Moh. Yamin Darsyah Multiyaningrum, Riska Nida Faoziatun Khusna Nida Faoziatun Khusna Ninu, Maria Febronia Nugrahanto, Rifqi Oktavia Sri Banowati Pandiriyan, Muhammad Tegar Permatasari, Shella Heidy Prizka Rismawati Arum Putra, Septian Malik Putri Ayu Firnanda Putri, Dinda Meyda Putri, Melfia Verahma Qonita Syalsabilla Handayani Rahma Nurmalita Rahmah, Alfidha Rahman, Budiono Rahmawati, Gita Ramadhan, Abimanyu Arya Rhendy K P Widiyanto Rizma Novinda Puteri Rochdi Wasono RR. Ella Evrita Hestiandari Safril Ahmadi Sanmas Safril Ahmadi Sanmas Salmaa Fauziah Sam'an, Muhammad Sanmas, Safril Ahmadi Sarah, Albertus Dion Septi Winda Utami Setiayani, Wiwik Shinta Amaria Shinta Amaria Soffi Amalia Nur Kholifah Syafira, Elsa Izza Syaharani, Nabbila Dyah Syahrani, Nabbila Dyah Syaifullah, Ahmad Reyhan Syifa Aulia Syifa Aulia Tari Fitri Ningsih Tiani Wahu Utami Tiani Wahyu Utami Watur, Annisa Cahyaningrum Widiyanti, Karin Dita Widiyanto, Rhendy K P Wiwit Putri Nur Izzaturrohmah Wulan Sari, Wulan Yan Nazala Bisoumi Yan Nazala Bisoumi Yuliardi, Fahrul Raditiar Yuni Nurkuntari