Claim Missing Document
Check
Articles

Found 10 Documents
Search

Uji Performa Algoritma Naïve Bayes untuk Prediksi Masa Studi Mahasiswa Irkham Widhi Saputro; Bety Wulan Sari
Creative Information Technology Journal Vol 6, No 1 (2019): Januari - Juni
Publisher : UNIVERSITAS AMIKOM YOGYAKARTA

Show Abstract | Download Original | Original Source | Check in Google Scholar | Full PDF (216.226 KB) | DOI: 10.24076/citec.2019v6i1.178

Abstract

Universitas AMIKOM Yogyakarta adalah salah satu perguruan tinggi yang memiliki ribuan mahasiswa baru khususnya pada prodi Informatika. Pada tahun 2012 tercatat ada 1009 mahasiswa baru, dan pada tahun 2013 juga tercatat ada sebanyak 859 mahasiswa baru. Namun sayangnya, dari sekian banyak mahasiswa hanya sekitar 50% saja yang dapat lulus dengan tepat waktu. Data tersebut untuk membuat sistem klasifikasi menggunakan teknik data mining dengan metode Naïve Bayes. Dataset yang akan digunakan sebanyak 300 data yang bersumber dari data alumni angkatan 2012, dan 2013 dengan masing-masing data sebanyak 150. Data yang diperoleh memiliki 144 mahasiswa dengan keterangan lulus tepat waktu, dan 156 mahasiswa dengan keterangan lulus tidak tepat waktu. Proses pengujian akan dilakukan menggunakan metode 10-Fold Cross Validation, dan Confusion Matrix. Hasil pengujian menunjukkan bahwa rata-rata performa dari model Naïve Bayes mempunyai nilai akurasi sebesar 68%, nilai precision sebesar 61.3%, nilai recall sebesar 65.3%, dan nilai f1-score sebesar 61%. Nilai performa dari model dapat dipengaruhi oleh dataset yang digunakan untuk pembuatan model.Kata Kunci — data mining, Naïve Bayes, K-Fold Cross Validation, Confusion MatrixAMIKOM Yogyakarta University is one of the colleges that has thousands of new students, especially in the Informatics study program. In 2012 there were 1009 new students, and in 2013 there were 859 new students. But unfortunately, of the many students only around 50% can graduate on time. The data is to make the classification system using data mining techniques with the Naïve Bayes method. The dataset will be used as much as 300 data sourced from alumni data of 2012, and 2013 with each data as much as 150. The data obtained has 144 students with information passed on time, and 156 students with graduation information not on time. The testing process will be carried out using the 10-Fold Cross Validation, and Confusion Matrix method. The test results show that the average performance of the Naïve Bayes model has an accuracy value of 68%, precision value is 61.3%, recall value is 65.3%, and f1-score is 61%. The performance value of the model can be influenced by the dataset used for modeling.Keywords — data mining, classification, Naïve Bayes, graduation time
Penerapan Konsep Gamification pada Pembelajaran Tenses Bahasa Inggris Berbasis Web Bety Wulan Sari; Ema Utami; Hanif Al Fatta
SISFOTENIKA Vol 5, No 2 (2015): SISFOTENIKA
Publisher : STMIK PONTIANAK

Show Abstract | Download Original | Original Source | Check in Google Scholar | Full PDF (566.268 KB) | DOI: 10.30700/jst.v5i2.87

Abstract

AbstrakTenses merupakan suatu bentuk kata kerja dalam tata bahasa pada bahasa Inggris yang berhubungan dengan waktu terjadinya suatu peristiwa. Dengan memahami dan menguasai tenses, maka bahasa Inggris akan terasa mudah karena tenses adalah dasar dari suatu pola kalimat bahasa Inggris. Terdapat 12 tenses yang sering digunakan dan seringkali membingungkan dan rumit bagi sebagian besar orang. Saat ini, e-learning semakin berkembang seiring dengan pesatnya kemajuan teknologi, sehingga siapapun yang membutuhkannya dapat mengakses darimanapun ia berada. Akan tetapi mayoritas e-learning saat ini tidak mampu menarik perhatian dan minat dari penggunanya. Konsep gamification akan terapkan ke dalam e-learning tenses bahasa Inggris agar pembelajaran lebih menarik dan menyenangkan.Perancangan gamified system ini akan menggunakan Marczewski’s Gamification Framework. Framework yang memiliki user types dengan kebutuhan pembelajaran dan pengembangan diri. Game mechanics pada user types ini seperti levels, challenges, dan rewards dapat mendukung pengguna gamified system dalam mencapai goals. Pengujian yang telah dilakukan menunjukkan keberhasilan dalam penerapan framework Marczewski terhadap kebutuhan fungsional pengguna. Kata Kunci : tenses, gamification, Marczewski’s Gamification Framework
PREDIKSI PEMBERIAN KELAYAKAN PINJAMAN DENGAN METODE FUZZY TSUKAMOTO Nurul Ajeng; Bety Wulan Sari; Donni Prabowo
Information System Journal Vol. 3 No. 1 (2020): Information System Journal (INFOS)
Publisher : Universitas Amikom Yogyakarta

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.24076/infosjournal.2020v3i1.215

Abstract

Sentra Gadai is a place to borrow in Yogyakarta. Every day giving loans to customers. In granting a loan the Senta Gadai has a condition that is a loan size of 50% of the collateral price. If the loan is more than 50%, the Sentra Gadai sometimes still hesitate to provide the loan. The Loan Eligibility Prediction System is used to help the Sentra Gadai in making decisions by providing alternative estimates in determining the feasibility of borrowing by the customer. This prediction system uses Tsukamoto's fuzzy method in estimating the feasibility of loans to customers by having several criteria such as the duration of the loan, the price of the guarantee and the condition of the goods. This prediction system is based on desktop because it is only used by the Sentra Gadai and not to the public with the Java programming language and database using phpMyAdmin. Keywords : Prediction System, Loan,Fuzzy Tsukamoto
NAIVE BAYES ALGORITHM IMPLEMENTATION TO DETECT HUMAN PERSONALITY DISORDERS Yoga Aditama Ika Nanda; Bety Wulan Sari
Jurnal Techno Nusa Mandiri Vol 17 No 1 (2020): Techno Nusa Mandiri : Journal of Computing and Information Technology Period of
Publisher : Lembaga Penelitian dan Pengabdian Pada Masyarakat

Show Abstract | Download Original | Original Source | Check in Google Scholar | Full PDF (991.672 KB) | DOI: 10.33480/techno.v17i1.1239

Abstract

We live in a society that still sees problems regarding one's soul and personality as taboo, even though mental health is as important as physical health. A personality disorder itself is a disorder that can be seen from behavior, mindset, and attitude, which brings difficulties to life. Based on this problem, this study applies the method of Naive Bayes classifier as early detection of human personality disorders. Using a data set of 130 correspondences from the AMIKOM university scope with the age limit of 18-25 years and identified personality disorders is a borderline type disorder. The data obtained was 94 with undiagnosed classes and 36 with undiagnosed classes, with the research variables in the form of questionnaire questions as many as 13 questions. The testing process is done with 10 fold and 5 fold cross-validation, and confusion matrix with the results in the form of accurate 10 folds superior with a value of 88.8% compared to 5 folds that is 88.2%, for precision 10 folds superior with 88.7%, but for 5 fold recall superior with 88.3%, while the final results of these two performances in F1-Score, produce the same value, which is 86.1%.
ANALISIS SIMULASI PREDIKSI CUSTOMER CHURN E-COMMERCE MENGGUNAKAN ALGORITMA RANDOM FOREST BERBASIS DATA SINTETIS Rayhan Bagoes Santoso; Bety Wulan Sari
JMBI UNSRAT (Jurnal Ilmiah Manajemen Bisnis dan Inovasi Universitas Sam Ratulangi). Vol 13 No 1 (2026): JMBI UNSRAT Volume 13 Nomor 1
Publisher : Magister Manajemen Program Pasca Sarjana Universitas Sam Ratulangi Manado

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.35794/jmbi.v13i1.67545

Abstract

Tingkat customer churn yang tinggi menjadi tantangan kritis bagi industri e-commerce di Indonesia dengan potensi kerugian mencapai miliaran rupiah per tahun. Penelitian ini bertujuan untuk mengimplementasikan sistem prediksi churn pelanggan sebagai proof-of-conceptmenggunakan algoritma machine learning. Mengingat keterbatasan akses data privat e-commerce, penelitian ini menggunakan pendekatan metodologis dengan dataset sintetis yang terdiri dari 1000 data pelanggan dan 9 fitur utama meliputi tenure, monthly spending, total transactions, support tickets, dan last purchase days. Tiga algoritma machine learning diimplementasikan yaitu Logistic Regression, Decision Tree, dan Random Forest untuk melakukan klasifikasi prediksi churn. Hasil penelitian menunjukkan bahwa Random Forest memberikan performa yang paling stabil dengan akurasi sebesar 87.5%, precision 86%, recall 87%, dan F1-score 86%. Penurunan performa dibandingkan model deterministik menunjukkan bahwa model diuji pada kondisi data yang lebih realistis dan tidak mengalami overfitting terhadap aturan generatif.
Analisis Perbandingan Prediksi Harga Rumah Dengan Random Forest, Gradient Boosting, dan XGBoost Bety Wulan Sari; Donni Prabowo
Intellect : Indonesian Journal of Learning and Technological Innovation Vol. 4 No. 1 (2025): Intellect : Indonesian Journal of Learning and Technological Innovation
Publisher : Yayasan Lembaga Studi Makwa

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.57255/intellect.v4i1.1385

Abstract

House price prediction poses a significant challenge in the property sector, especially in the Yogyakarta region, which exhibits a wide range of price variations. This study aims to compare the performance of three regression algorithms such as Random Forest, Gradient Boosting, and XGBoost, in building predictive models based on features such as land area, building area, number of bedrooms, bathrooms, and garage availability. The dataset analyzed consists of 1,642 entries, with house prices ranging from IDR 7 million to IDR 4.37 billion, an average price of IDR 1.14 billion, and a mode of IDR 775 million. Model evaluation was conducted using Mean Squared Error (MSE) and the coefficient of determination (R²), where XGBoost achieved the best performance with an MSE of 1.56 × 10¹⁴ IDR², an R² of 0.7746, and a Root Mean Squared Error (RMSE) of approximately IDR 12.5 million. These results indicate that XGBoost outperforms the other two models in handling complex tabular data and provides more accurate predictions. The predictive model has practical potential to be utilized by property developers, real estate agents, and local governments as a decision-support tool for price estimation, market evaluation, and data-driven urban planning. These findings highlight that selecting the appropriate algorithm can significantly enhance the quality of house price prediction. Abstrak Prediksi harga rumah menjadi tantangan penting dalam bidang properti, khususnya di wilayah Yogyakarta yang memiliki variasi harga cukup ekstrem. Penelitian ini bertujuan untuk membandingkan performa tiga algoritma regresi yaitu Random Forest, Gradient Boosting, dan XGBoost digunakan untuk membangun model prediksi harga rumah berdasarkan fitur seperti luas tanah, luas bangunan, jumlah kamar tidur, kamar mandi, dan garasi. Data yang dianalisis mencakup 1.642 entri dengan harga rumah berkisar antara Rp 7 juta hingga Rp 4,37 miliar, harga rata-rata sebesar Rp 1,14 miliar, dan modus Rp 775 juta. Evaluasi model dilakukan menggunakan metrik Mean Squared Error (MSE) dan koefisien determinasi (R²), di mana XGBoost menghasilkan performa terbaik dengan MSE sebesar 1,56 × 10¹⁴ rupiah², R² sebesar 0,7746, dan Root Mean Squared Error (RMSE) sekitar 12,5 juta rupiah. Hasil ini menunjukkan bahwa XGBoost lebih unggul dalam menangani data tabular kompleks dan memiliki akurasi prediksi yang lebih baik dibanding dua model lainnya. Model prediktif ini berpotensi digunakan oleh pengembang properti, agen real estate, maupun pemerintah daerah sebagai alat bantu dalam penetapan harga, evaluasi pasar, dan perencanaan tata ruang yang berbasis data. Temuan ini memberikan gambaran bahwa pemilihan algoritma yang tepat dapat meningkatkan kualitas prediksi harga properti.
Cross-Dataset Evaluation of Boosting Models for Hypertension Prediction Bety Wulan Sari; Dewi Ayu Murtiningsih; Donni Prabowo; Yoga Pristyanto; Ika Nur Fajri; Ike Verawati
Jurnal RESTI (Rekayasa Sistem dan Teknologi Informasi) Vol 10 No 4 (2026): August 2026 (in progress)
Publisher : Ikatan Ahli Informatika Indonesia (IAII)

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.29207/resti.v10i4.7620

Abstract

Hypertension remains a significant global risk factor for cardiovascular disease and related mortality, necessitating reliable early risk prediction models. Although boosting algorithms have demonstrated strong performance in structured medical data, limited studies have examined their consistency across heterogeneous datasets. This study aims to evaluate the cross-dataset performance and stability of three boosting models, such as XGBoost, LightGBM, and CatBoost, for hypertension prediction under multiple train–test split ratios. Two independent structured datasets were analyzed using 60:40, 70:30, 80:20, and 90:10 splits. To identify the optimal hyperparameters, grid search was performed using repeated stratified 5-fold cross-validation with three repetitions. Model effectiveness was measured using the evaluation metrics of accuracy, precision, recall, F1-score, and AUC. Results show that Dataset 1 gained consistently high predictive performance (accuracy > 0.98; AUC ≈ 1.00), indicating strong and well-separated predictive signals, whereas Dataset 2 demonstrated substantially lower discriminative ability (accuracy ≈ 0.71–0.72; AUC ≈ 0.50), suggesting limited predictive structure. Across both datasets, CatBoost consistently obtained the highest accuracy, particularly at the 90:10 split ratio. These findings demonstrate that dataset characteristics critically determine model effectiveness and that among the evaluated boosting algorithms, CatBoost delivered the strongest overall predictive performance.
Provincial Clustering using GARCH-based Chili Price Volatility Features Yogata Rama Guninta; Bety Wulan Sari; Yoga Pristyanto
SISTEMASI Vol 15, No 7 (2026): Sistemasi: Jurnal Sistem Informasi
Publisher : Universitas Islam Indragiri

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.32520/stmsi.v15i7.6521

Abstract

Bird's eye chili is a strategic food commodity in Indonesia whose prices are highly susceptible to interregional fluctuations due to differences in distribution systems and supply chain conditions. These fluctuations often occur over short time horizons, making analyses based on monthly or annual data less capable of capturing short-term price spikes that have the greatest impact on consumers' purchasing power and price stabilization policies. Previous studies have generally clustered regions based on nominal or average prices, which do not adequately represent the dynamics of daily price movements. This study aims to cluster Indonesian provinces according to the volatility characteristics of bird's eye chili prices by proposing a clustering approach that utilizes GARCH(1,1)-based conditional volatility features to represent daily price dynamics, combined with the K-Means algorithm for cluster formation. Daily bird's eye chili price data from 34 provinces covering the period from April 2024 to April 2026 were obtained from the National Strategic Food Price Information Center (PIHPS). The price data were transformed into daily returns and modeled using GARCH(1,1) to estimate the average conditional volatility of each province, which was subsequently used as the clustering feature. The optimal number of clusters was determined using the Elbow Method and the Silhouette Score. The evaluation results identified five as the optimal number of clusters, achieving a silhouette score of 0.65 and classifying the 34 provinces into five volatility categories: very high, high, moderate, low, and very low. The findings reveal that provinces with higher price levels do not necessarily belong to the highest volatility cluster, indicating that a volatility-based approach provides additional insights into price dynamics beyond those captured by nominal price-based clustering. These results can support regional food price volatility monitoring and serve as a reference for developing data-driven decision support systems for food price management.
KLASIFIKASI VARIETAS BIBIT DURIAN MENGGUNAKAN RESNET50: PENDEKATAN DEEP LEARNING UNTUK PERTANIAN DIGITAL Janottama Kalam Putra Sucipto; Bety Wulan Sari
Journal of Information System Management (JOISM) Vol. 7 No. 2 (2026): Januari
Publisher : Universitas Amikom Yogyakarta

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.24076/joism.2026v7i2.2479

Abstract

Identifikasi varietas bibit durian secara akurat pada fase pembibitan sangat krusial untuk mencegah kerugian ekonomi akibat kesalahan pemilihan varietas unggul. Namun, identifikasi manual berbasis visual memiliki kelemahan pada subjektivitas dan tingkat kesalahan manusia yang tinggi. Penelitian ini mengusulkan model klasifikasi otomatis untuk empat varietas durian populer (Bawor, Duri Hitam, Musang King, dan Super Tembaga) menggunakan arsitektur Deep Residual Network (ResNet50). Peningkatan akurasi dilakukan melalui integrasi teknik prapemrosesan background removal berbasis ambang batas untuk mereduksi noise latar belakang dan penerapan strategi fine-tuning pada lapisan fully connected. Selain itu, optimasi hyperparameter dilakukan secara sistematis untuk menentukan learning rate dan batch size optimal. Hasil eksperimen menunjukkan bahwa model yang diusulkan mencapai performa superior dengan akurasi klasifikasi sebesar 96% dan stabilitas nilai loss pada rentang 0.15–0.20. Hasil ini membuktikan bahwa pendekatan deep learning dengan optimasi prapemrosesan mampu memberikan solusi identifikasi yang lebih objektif dan presisi dibandingkan metode konvensional. Penelitian ini berkontribusi pada pengembangan sistem pertanian digital yang cerdas dalam mendukung standarisasi kualitas bibit durian.
Enhanced Predictive Modeling for Non-Invasive Liver Disease Diagnosis Donni Prabowo; Bety Wulan Sari; Yoga Pristyanto; Afrig Aminuddin
Jurnal RESTI (Rekayasa Sistem dan Teknologi Informasi) Vol 9 No 4 (2025): August 2025
Publisher : Ikatan Ahli Informatika Indonesia (IAII)

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.29207/resti.v9i4.6449

Abstract

Liver diseases (e.g. cirrhosis, hepatitis, and fatty liver disease) are globally one of the leading causes of mortality and are typically diagnosed in advanced stages due to vague symptoms and the difficulty involved in existing diagnostic techniques (e.g. biopsies). To optimize the early diagnosis of liver disease, this study proposes an enhanced, non-invasive approach using machine learning techniques. The research is enriched with a full pipeline, from exploratory data analysis and imputation of the dataset, treatment of the outlier, encoding of labels and scaling using ILPD (Indian Liver Patient Dataset). The classification models compared were RandomForest, XGBoost, LGBM, and CatBoost. The CatBoost algorithm fine-tuned with RandomizedSearchCV showed the highest performance with a test accuracy of 93%. The performance was again better than any already published methods showing that advanced ensembling and hyperparameter optimization worked. The proposed model is suitable for incorporation into clinical decision support systems and provides reliable and accurate diagnostic assistance. In addition to its high accuracy, the model is robust for missing and categorical data, which is a challenge in any real-world clinical scenario. These findings add to the growing body of evidence supporting AI-based medical diagnostics and suggest that CatBoost is a highly promising tool for facilitating timely screening and diagnosis of liver disease. Furthermore, the study stresses the need for thorough preprocessing and cross-validation, which serve to reduce biases that are present in widely applied datasets. Ongoing future efforts may involve the integration of multi-source data and implementation of explainable AI techniques to allow for wider clinical trust and use.