Claim Missing Document
Check
Articles

Analisis Komparatif Jarak Euclidean, Manhattan, Canberra, Chebyshev, Cosine pada K-Means untuk Evaluasi Kepuasan Masyarakat Fakhri Fakhri; Iis afrianty; Elvia Budianita; Fadhilah Syafria; Siska Kurnia Gusti; Salmiyati Salmiyati
TIN: Terapan Informatika Nusantara Vol 7 No 1 (2026): June 2026
Publisher : Forum Kerjasama Pendidikan Tinggi (FKPT)

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47065/tin.v7i1.10086

Abstract

The selection of distance metrics in the K-Means Clustering algorithm can affect the quality of clustering results, particularly on public satisfaction data measured using a Likert scale. This study aims to compare the performance of five distance metrics, namely Euclidean Distance, Manhattan Distance, Canberra Distance, Chebyshev Distance, and Cosine Similarity, in clustering the level of public satisfaction toward public services. The research data were obtained from 533 respondents who used the services of the Mal Pelayanan Publik (MPP) Pekanbaru through a questionnaire consisting of 23 questions based on the SERVQUAL dimensions and the Community Satisfaction Survey indicators in accordance with PERMEN PAN-RB Number 14 of 2017. After the data cleaning process, one duplicate record was removed, resulting in 532 respondent records used in the analysis stage. The number of clusters was determined using the Elbow Method, while cluster quality was evaluated using the Davies-Bouldin Index (DBI) and Silhouette Score. The results show that Manhattan Distance with k=2 produced the lowest DBI value of 0.8144, whereas Euclidean Distance with k=3 produced the highest Silhouette Score of 0.5088. The clustering results formed groups of respondents with different satisfaction levels, namely Dissatisfied, Satisfied, and Very Satisfied. This study contributes an evaluative comparison of five distance metrics in the K-Means algorithm using two evaluation approaches simultaneously, namely the Davies-Bouldin Index and Silhouette Score, on public satisfaction data based on a Likert scale. The results indicate that the performance of distance metrics may differ depending on the evaluation method used, therefore the selection of distance metrics should consider the characteristics of the data and the objectives of the analysis.The difference in evaluation results indicates that DBI and Silhouette Score assess clustering quality from different aspects. Based on the findings, Manhattan Distance and Euclidean Distance demonstrated better performance compared to other distance metrics on the dataset used, and can therefore be considered in the analysis of public satisfaction toward public services.
Information Gain and Random Forest for Sex Classification Based on Craniometric Measurements Nabilla Alya Firana; Iis Afrianty; Novriyanto Novriyanto; Febi Yanto
Journal of Artificial Intelligence and Software Engineering Vol 6, No 2 (2026): Juni (OnProgress)
Publisher : Politeknik Negeri Lhokseumawe

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.30811/jaise.v6i2.9409

Abstract

Sex identification from human skulls is a crucial aspect of forensic anthropology; however, traditional methods still face limitations such as subjective assessment and inter-population variation. This study proposes the application of Information Gain as a feature selection technique and Random Forest as a classification algorithm for sex determination based on craniometric data. The dataset used is the Howells dataset consisting of 2,524 samples with 83 skull measurement features. Feature selection using Information Gain was performed with threshold values of 0.01, 0.05, and 0.09, followed by additional testing across a threshold range of 0.01 to 0.09. Model evaluation was conducted using 10-Fold Cross Validation with default Random Forest parameters. The results show that a threshold of 0.02 produced 57 selected features from the original 83, achieving the best performance with an accuracy of 87.40%, precision of 87.53%, recall of 87.40%, and F1-score of 87.41%. These results outperform the baseline model without feature selection, which achieved an accuracy of 86.57%. This study demonstrates that Information Gain feature selection can reduce data dimensionality by 31.3% while simultaneously improving sex classification performance based on craniometric data.
Application of Information Gain Feature Selection and SMOTE in XGBoost Algorithm for Asthma Disease Classification Fioni Nikmatul Fajar; Fitri Insani; Suwanto Sanjaya; Iis Afrianty
Journal of Artificial Intelligence and Software Engineering Vol 6, No 2 (2026): Juni (OnProgress)
Publisher : Politeknik Negeri Lhokseumawe

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.30811/jaise.v6i2.9384

Abstract

Asma merupakan salah satu penyakit kronis pada sistem pernapasan yang prevalensinya terus meningkat dan memerlukan deteksi dini untuk mencegah komplikasi serius. Salah satu tantangan dalam klasifikasi asma menggunakan machine learning adalah ketidakseimbangan kelas yang menyebabkan model cenderung memprediksi kelas mayoritas sehingga kemampuan mendeteksi kasus asma menjadi rendah. Penelitian ini mengusulkan penerapan SMOTE dan seleksi fitur Information Gain dalam algoritma XGBoost untuk mengatasi permasalahan tersebut. Dataset yang digunakan terdiri dari 2.392 data dengan 28 atribut, di mana tahapan penelitian meliputi preprocessing, seleksi fitur menggunakan Information Gain yang mengurangi fitur menjadi 22 fitur, penyeimbangan data menggunakan SMOTE, pembagian data dengan rasio 90:10, 80:20, dan 70:30, serta klasifikasi menggunakan XGBoost. Pengujian dilakukan terhadap empat skenario pendekatan untuk membandingkan kontribusi setiap metode yang diterapkan. Evaluasi dilakukan menggunakan data uji seimbang dan data uji asli dengan metrik akurasi, presisi, recall, dan F1-score. Hasil penelitian menunjukkan bahwa skenario terbaik diperoleh pada kombinasi Information Gain + SMOTE + XGBoost dengan rasio 90:10 pada data uji seimbang, menghasilkan akurasi 75%, presisi 87,5%, recall 58,33%, dan F1-score 70%. Hasil tersebut menunjukkan bahwa kombinasi seleksi fitur dan penyeimbangan data mampu meningkatkan kemampuan model dalam mendeteksi penyakit asma.
GAMBARAN KARAKTERISTIK IBU POST SECTIO CESAREA TERKAIT PENYEMBUHAN LUKA Ekawati Saputri, Saputri; Iis Afrianty; Evodius Nasus
Jurnal Ilmu Kesehatan Abdurrab Vol. 1 No. 4 (2023): Vol 1 No 4 Desember 2023
Publisher : LPPM Universitas Abdurrab

Show Abstract | Download Original | Original Source | Check in Google Scholar

Abstract

Sectio caesarea is a delivery method that is carried out by making an open incision in the uterine wall, causing wounds in the abdominal area. Globally, around 21% of caesarean section deliveries occur. In Indonesia, caesarean section delivery is around 17.6%. This study aims to determine the characteristics of post-cesarean section mothers regarding wound healing. This research is descriptive quantitative research with a Secondary Data Analysis (ADS) approach. The total sample was 136 using purposive sampling technique. The results of this study show that the characteristics of mothers post cesarean section are that most of them are aged 20-35 years (76.5%) with secondary education level (46.3%), multiparous (69.1%), and have no history of CS (56, 6%) and did not suffer from anemia (61.8%). Almost all of the wounds experienced by mothers after caesarean section were dry wounds (99.3%). The post caesarean section wound is healing well.
Penerapan Metode ADASYN Dalam Mengatasi Imbalanced Data Untuk Klasifikasi Penyakit Stroke Menggunakan Support Vector Machine Alwaliyanto Alwaliyanto; Siska Kurnia Gusti; Iis Afrianty; Fadhilah Syafria
Bulletin of Computer Science Research Vol. 5 No. 4 (2025): June 2025
Publisher : Forum Kerjasama Pendidikan Tinggi (FKPT)

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47065/bulletincsr.v5i4.612

Abstract

Stroke is one of the leading causes of death and disability worldwide, making it essential to develop classification models that can assist in early and accurate diagnosis. This study aims to implement the Support Vector Machine (SVM) algorithm with three types of kernels linear, polynomial, and Radial Basis Function (RBF) to classify stroke disease data. The Adaptive Synthetic Sampling (ADASYN) method is employed to address the class imbalance problem, while model training and evaluation are carried out using 5-Fold Cross-Validation to ensure stable and reliable results. The findings indicate that ADASYN successfully improves the model’s sensitivity to stroke cases (the minority class), as reflected by an increase in recall and F1-score, despite a slight decrease in overall accuracy a common trade-off in handling imbalanced data. The linear kernel (after ADASYN) achieved the best performance after imbalance handling, with an average AUC-ROC of 0.8333, recall of 0.7827, and F1-score of 0.2181 for the stroke class. Although the F1-score remains relatively low, it improved compared to the pre-ADASYN results, indicating better detection of stroke cases. The implementation was conducted using Google Colab, which also contributed to efficient data processing and visualization. Overall, the results demonstrate that the combination of SVM and ADASYN is effective in enhancing the model’s sensitivity to minority classes and is well-suited for medical data classification tasks, particularly in the early diagnosis of stroke using machine learning approaches.
Penerapan Information Gain Untuk Seleksi Fitur Pada Klasifikasi Jenis Kelamin Tulang Tengkorak Menggunakan Backpropagation Nada Tsawaabul Khair; Iis Afrianty; Fadhilah Syafria; Elvia Budianita; Siska Kurnia Gusti
Bulletin of Computer Science Research Vol. 5 No. 4 (2025): June 2025
Publisher : Forum Kerjasama Pendidikan Tinggi (FKPT)

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47065/bulletincsr.v5i4.637

Abstract

Forensic anthropology and skull analysis play a crucial role in the biological identification of individuals, including sex determination. This study aims to improve the accuracy of gender classification based on skull structure by combining the Information Gain feature selection method with the Backpropagation algorithm. The dataset used is the craniometric data compiled by William W. Howells, consisting of 2,524 samples with 85 measurement features. The preprocessing stage includes data selection, data cleaning, and normalization. Feature selection was conducted using the Information Gain method with three threshold values: 0.01, 0.05, and 0.1, resulting in 79, 46, and 38 selected features, respectively. The model was evaluated using the K-Fold Cross Validation method with K=10 and K=20. The highest accuracy of 93.91% was achieved at the 0.01 threshold using the Backpropagation architecture [79:119:1], a learning rate of 0.01, and K=20. These results demonstrate that feature selection using Information Gain enhances the performance of the Backpropagation model by eliminating irrelevant features and minimizing the risk of overfitting.
Perbandingan Teknik Penyeimbang Kelas Pada Multi-Layer Perceptron (MLP) Berbasis Backpropagation Untuk Klasifikasi Diabetes Mellitus Robby Azhar; Siska Kurnia Gusti; Iis Afrianty; Elvia Budianita
Bulletin of Computer Science Research Vol. 5 No. 6 (2025): October 2025
Publisher : Forum Kerjasama Pendidikan Tinggi (FKPT)

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47065/bulletincsr.v5i6.804

Abstract

Diabetes Mellitus (DM) is a chronic disease that can lead to serious complications if not detected early; therefore, early diagnosis is highly important. One of the methods that can be applied for early diagnosis is the classification technique in data mining. However, the classification process often faces challenges due to class imbalance, which can reduce model performance. This study aims to analyze the effect of class balancing techniques on the performance of the Backpropagation Neural Network (BPNN) in classifying DM cases. BPNN is a form of Multi-Layer Perceptron (MLP) with a simple structure and the ability to solve complex problems with good accuracy. The dataset used in this study is the Pima Indians Diabetes Dataset, consisting of 768 instances, including 500 non-diabetic and 268 diabetic cases. The research was conducted using three scenarios: without balancing, Synthetic Minority Over-sampling Technique (SMOTE), and Random Under Sampling (RUS). The BPNN model was designed with two architectural variations (one hidden layer and two hidden layers), three learning rate values (0.1, 0.01, and 0.001), and a varying number of neurons. The dataset was divided using the 10-Fold Cross Validation technique. The results show that applying SMOTE achieved the best performance, with an average accuracy of 90.89%, precision of 91.22%, recall of 90.89%, and F1-score of 90.89% on the BPNN architecture with one hidden layer. Furthermore, the single hidden layer architecture proved more stable than the two hidden layers, especially when the dataset size decreased due to RUS. Therefore, the combination of SMOTE and BPNN with one hidden layer provides better performance in classifying Diabetes Mellitus cases.
Penerapan Seleksi Fitur Information Gain dan Metode Backpropagation Neural Network Untuk Klasifikasi Atrisi Karyawan Dinyah Fithara; Elvia Budianita; Iis Afrianty; Siska Kurnia Gusti
Bulletin of Computer Science Research Vol. 6 No. 1 (2025): December 2025
Publisher : Forum Kerjasama Pendidikan Tinggi (FKPT)

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47065/bulletincsr.v6i1.922

Abstract

Employee attrition management is a critical challenge for organizations as it involves costs, time, and the risk of decision-making errors. This problem requires a data-driven business strategy to achieve more accurate predictions of employees who are potentially at risk of termination. This study applies the Information Gain feature selection method and the Backpropagation Neural Network (BPNN) algorithm in the employee attrition classification process with the aim of increasing the accuracy and efficiency of the prediction model. BPNN is chosen due to its simpler architecture, faster training time, and greater stability for small to medium sized datasets.  With the assistance of Information Gain feature selection, BPNN is able to achieve optimal performance without requiring a complex architecture. The dataset used consist of 35 attributes and 1.470 employee records covering various factor such as age, income level, and employment status. The research stages include feature selection based on information gain values with specific thresholds, data partitioning using k-fold cross validation, and model training using BPNN with variations of learning rates and hidden neuron counts. The results show that the combination of Information Gain and BPNN improves classification accuracy compared to models without feature selection, achieving the highest average accuracy of 87.28% when using 25 selected attributes, with a BPNN configuration of learning rate 0.001, 35 hidden neurons, and 50 epochs. The attributes with the highest Information Gain score include JobLevel, OverTime, MaritalStatus, and MonthlyIncome. This study demonstrates that the proposed approach successfully enhances the prediction performance of employee attrition and can serve as a foundation for developing data-driven models that support employee retention efforts.
Sistem Prediksi Produksi Kelapa Sawit Berbasis Gradio Menggunakan Algoritma Regresi Linear Berganda Irfan Jamal Matondang; Elvia Budianita; Fadhilah Syafria; Iis Afrianty
Bulletin of Computer Science Research Vol. 6 No. 2 (2026): February 2026
Publisher : Forum Kerjasama Pendidikan Tinggi (FKPT)

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47065/bulletincsr.v6i2.994

Abstract

The instability of oil palm production often leads to discrepancies between production targets and actual outputs, thereby necessitating an accurate prediction model to support operational planning. This study aims to develop an oil palm production prediction model and to identify the most influential variables affecting production outcomes as a basis for data-driven decision-making. The model was developed using the Multiple Linear Regression method based on historical data from 2020–2024, consisting of 60 monthly observations with variables including number of trees, land area, rainfall, number of fruit bunches, and plant age. The research stages included data preprocessing, variable selection through testing several feature combinations, model development, and performance evaluation using the coefficient of determination (R²), Mean Absolute Error (MAE), Mean Squared Error (MSE), and Root Mean Square Error (RMSE). The results indicate that the combination of number of trees, land area, number of fruit bunches, and plant age produced the best performance, with an R² value of 0.85 on the training data and 0.81 on the testing data. The MAE values were 125,307 kg and 176,984 kg, the MSE values were 28,870,838,455 kg² and 52,809,954,662 kg², and the RMSE values were 169,914 kg and 229,804 kg, respectively. Based on the regression coefficients, the number of fruit bunches was identified as the most dominant variable, with a coefficient value of 637,720 kg. The model was subsequently implemented using the Python Gradio library in the form of an interactive interface to support production planning effectiveness and minimize the risk of inaccurate decision-making in oil palm plantation management.
Analisis Penerapan Flexmatch Pada Semi-Supervised Deep Learning Untuk Deteksi Penyakit Paru-Paru Berbasis Citra Chest X-Ray M. Aufaa Rahman; Benny Sukma Negara; Muhammad Irsyad; Lestari Handayani; Iis Afrianty
Jurnal Pengembangan Teknologi Informasi dan Komunikasi (JUPTIK) Vol. 4 No. 1 (2026): JURNAL PENGEMBANGAN TEKNOLOGI INFORMASI DAN KOMUNIAKSI (JUPTIK)
Publisher : Universitas Muhammadiyah Muara Bungo

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.52060/juptik.v4i1.4439

Abstract

Penyakit paru-paru seperti COVID-19 dan Pneumonia merupakan penyebab utama morbiditas dan mortalitas di dunia, sehingga diperlukan metode deteksi dini yang akurat dan efisien. Penelitian ini bertujuan menganalisis penerapan FlexMatch pada skema Semi-Supervised Deep Learning untuk klasifikasi penyakit paru-paru berbasis citra Chest X-ray (CXR), dengan memanfaatkan data berlabel dan tidak berlabel melalui mekanisme Curriculum Pseudo Labeling dan class-adaptive thresholding, serta DenseNet-169 sebagai ekstraktor fitur utama. Dataset yang digunakan terdiri dari 3.000 citra CXR yang mencakup tiga kelas, yaitu COVID-19, pneumonia, dan normal. Tahapan penelitian meliputi preprocessing data, augmentasi citra, pembagian data, pelatihan model, serta evaluasi menggunakan accuracy, precision, recall, F1-score, dan Grad-CAM. Hasil penelitian menunjukkan bahwa model mencapai akurasi validasi sebesar 96,65% dengan F1-score masing-masing sebesar 99,02% untuk COVID-19, 95,74% untuk Normal, dan 94,48% untuk Pneumonia. Visualisasi Grad-CAM membuktikan bahwa model mampu memfokuskan perhatian pada area paru yang relevan secara klinis, sehingga FlexMatch terbukti efektif meningkatkan performa klasifikasi pada kondisi data berlabel terbatas dan berpotensi mendukung sistem diagnosis penyakit paru-paru berbasis kecerdasan buatan.