Claim Missing Document
Check
Articles

Found 39 Documents
Search

Dialect Classification of the Javanese Language Using the K-Nearest Neighbor Filby, Brilliant; Pujianto, Utomo; Hammad, Jehad A. H.; Wibawa, Aji Prasetya
Journal of Information Technology and Cyber Security Vol. 2 No. 2 (2024): July
Publisher : Department of Information Systems and Technology, Faculty of Intelligent Electrical and Informatics Technology, Universitas 17 Agustus 1945 Surabaya

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.30996/jitcs.12213

Abstract

Indonesia is rich in ethnic and cultural diversity, each reflected in its unique linguistic characteristics. One way to preserve the Javanese language is by conducting research on its dialects. This study aims to classify three main dialects in Java Island—East Java, Central Java, and West Java—using text data from online sources. The classification process includes preprocessing (tokenizing, case folding, and word weighting), data balancing with the Synthetic Minority Oversampling Technique (SMOTE), and classification using the K-Nearest Neighbor (K-NN) algorithm. This study highlights the importance of dialect recognition in supporting the preservation of the Javanese language and the development of linguistic technology applications. Testing using 10-fold cross-validation showed the best performance at , with an accuracy of 94.05%, precision of 95.83%, and recall of 94.44%. These findings significantly support computational linguistics research and the preservation of regional languages.
PERBANDINGAN METODE NAÏVE BAYES DAN C4.5 UNTUK MEMPREDIKSI MORTALITAS PADA PETERNAKAN AYAM BROILER Baihaqi, Dimas Imam; Handayani, Anik Nur; Pujianto, Utomo
Simetris: Jurnal Teknik Mesin, Elektro dan Ilmu Komputer Vol 10, No 1 (2019): JURNAL SIMETRIS VOLUME 10 NO 1 TAHUN 2019
Publisher : Fakultas Teknik Universitas Muria Kudus

Show Abstract | Download Original | Original Source | Check in Google Scholar | Full PDF (185.135 KB) | DOI: 10.24176/simet.v10i1.2846

Abstract

Ayam broiler adalah jenis ternak yang paling cepat untuk dipanen. Namun dalam berternak ayam broiler pasti banyak masalah yang dihadapi contohnya adalah tingkat kematian. Untuk menekan kerugian, para peternak sebaiknya memperhatikan faktor-faktor apa saja yang menyebabkan kematian ayam tersebut. Beberapa penelitian yang meneliti tentang ayam broiler menggunakan metode percobaan dan RAL. Namun masih belum ada yang meneliti mortalitas ayam broiler menggunkan komputasi. Untuk mengetahui metode mana yang lebih baik untuk memprediksi mortalitas pada peternakan ayam broiler dilakukan penelitian perbandingan metode Naïve Bayes dan C4.5. Hasil dari perbandingan akan dievaluasi menggunakan confution matrix. Hasil dari pengujian data menggunakan confution matrix menghasilkan nilai akurasi dari metode C4.5 lebih besar dari pada metode Naïve Bayes. Nilai akurasi dari metode C4.5 adalah 93% dan nilai akurasi dari metode Naïve Bayes adalah 88.66%.
Journal Unique Visitors Forecasting Based on Multivariate Attributes Using CNN Dewandra, Aderyan Reynaldi Fahrezza; Wibawa, Aji Prasetya; Pujianto, Utomo; Utama, Agung Bella Putra; Nafalski, Andrew
International Journal of Artificial Intelligence Research Vol 6, No 2 (2022): Desember 2022
Publisher : Universitas Dharma Wacana

Show Abstract | Download Original | Original Source | Check in Google Scholar | Full PDF (379.839 KB) | DOI: 10.29099/ijair.v6i1.274

Abstract

Forecasting is needed in various problems, one of which is forecasting electronic journals unique visitors. Although forecasting cannot produce very accurate predictions, using the proper method can reduce forecasting errors. In this research, forecasting is done using the Deep Learning method, which is often used to process two-dimensional data, namely convolutional neural network (CNN). One-dimensional CNN comes with 1D feature extraction suitable for forecasting 1D time-series problems. This study aims to determine the best architecture and increase the number of hidden layers and neurons on CNN forecasting results. In various architectural scenarios, CNN performance was measured using the root mean squared error (RMSE). Based on the study results, the best results were obtained with an RMSE value of 2.314 using an architecture of 2 hidden layers and 64 neurons in Model 1. Meanwhile, the significant effect of increasing the number of hidden layers on the RMSE value was only found in Model 1 using 64 or 256 neurons.
Automatic Classification of Artificial Intelligence Generated Question Difficulty Levels: Klasifikasi Otomatis Tingkat Kesulitan Soal Hasil Kecerdasan Buatan Najahah, Vina; Pujianto, Utomo
Indonesian Journal of Innovation Studies Vol. 27 No. 1 (2026): January
Publisher : Universitas Muhammadiyah Sidoarjo

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.21070/ijins.v27i1.1880

Abstract

General Background: Determining question difficulty is a fundamental requirement in educational assessment to support valid evaluation and systematic question curation. Specific Background: The increasing use of artificial intelligence for automatic question generation produces large volumes of linguistically diverse items, making manual difficulty labeling time-consuming and subjective. Knowledge Gap: Despite extensive research on text-based difficulty prediction, lightweight and reproducible pipelines for multi-level difficulty classification of AI-generated questions remain limited. Aims: This study aims to develop and evaluate an automatic classification pipeline for three difficulty levels of AI-generated multiple-choice questions using TF-IDF text representation and a Random Forest classifier. Results: The proposed pipeline achieved a test accuracy of 70.98%, exceeding the random guessing baseline, with the highest F1-score observed in the easy class (78.45%) and the lowest in the medium class (65.32%), indicating greater ambiguity in intermediate difficulty questions. Novelty: This study presents a reproducible and interpretable classification workflow specifically applied to expert-labeled AI-generated questions with high inter-rater reliability. Implications: The findings support the use of lexical feature–based classification as an initial pre-curation and difficulty filtering tool in AI-assisted educational assessment systems. Highlights • The classification pipeline distinguishes three difficulty levels using only textual features• Medium difficulty questions exhibit the highest classification ambiguity• Lexical patterns contribute consistently to difficulty level separation Keywords Question Difficulty Classification; AI Generated Questions; TF-IDF Representation; Random Forest Classifier; Educational Assessment
Generative Artificial Intelligence Label Reliability in Programming Assessment: Reliabilitas Label Kecerdasan Buatan Generatif pada Asesmen Algoritma Pemrograman Khairunnisa, Raissa Araminta; Pujianto, Utomo
Indonesian Journal of Innovation Studies Vol. 27 No. 1 (2026): January
Publisher : Universitas Muhammadiyah Sidoarjo

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.21070/ijins.v27i1.1881

Abstract

General Background: The integration of Generative AI in educational assessment enables rapid construction of large-scale question banks, particularly in programming education, yet raises concerns regarding content validity. Specific Background: In algorithm and programming domains, Generative AI models frequently assign Higher Order Thinking Skills and Lower Order Thinking Skills labels automatically, creating potential discrepancies with Bloom’s Taxonomy classifications. Knowledge Gap: Empirical evidence validating the reliability of AI-generated cognitive labels and comparing statistical and transformer-based classification methods on small, domain-specific Indonesian datasets remains limited. Aims: This study aims to audit the reliability of cognitive labels generated by the Gemini model through expert validation and to compare TF-IDF–SVM and IndoBERT–SVM classifiers under class-imbalanced conditions. Results: Expert validation revealed substantial mislabeling, with a claimed balanced dataset becoming skewed toward LOTS. Classification experiments using five-fold cross-validation showed that TF-IDF–SVM achieved a slightly higher macro F1-score than IndoBERT–SVM. Novelty: The study demonstrates that simple lexical representations with stemming can outperform transformer-based embeddings when data are limited and domain-specific. Implications: These findings emphasize the necessity of human validation in AI-generated assessments and support the use of lightweight statistical text classification for automated cognitive level evaluation in constrained educational contexts. Highlights • Generative AI cognitive labels showed substantial inconsistency after expert validation• Lexical feature representation yielded higher macro-level classification balance• Human-in-the-loop validation remained essential for programming assessment datasets Keywords HOTS; LOTS; Generative AI; Text Classification; TF-IDF
LSTM-based Multivariate Time-Series Analysis: A Case of Journal Visitors Forecasting Saputra, Anggie Wahyu; Wibawa, Aji Prasetya; Pujianto, Utomo; Putra Utama, Agung Bella; Nafalski, Andrew
ILKOM Jurnal Ilmiah Vol 14, No 1 (2022)
Publisher : Prodi Teknik Informatika FIK Universitas Muslim Indonesia

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.33096/ilkom.v14i1.1106.57-62

Abstract

Forecasting is the process of predicting something in the future based on previous patterns. Forecasting will never be 100% accurate because the future has a problem of uncertainty. However, using the right method can make forecasting have a low error rate value to provide a good forecast for the future. This study aims to determine the effect of increasing the number of hidden layers and neurons on the performance of the long short-term memory (LSTM) forecasting method. LSTM performance measurement is done by root mean square error (RMSE) in various architectural scenarios. The LSTM algorithm is considered capable of handling long-term dependencies on its input and can predict data for a relatively long time. Based on research conducted from all models, the best results were obtained with an RMSE value of 0.699 obtained in model 1 with the number of hidden layers 2 and 64 neurons. Adding the number of hidden layers can significantly affect the RMSE results using neurons 16 and 32 in Model 1.
Rekomendasi Pemetaan Keahlian Siswa Terhadap Spesifikasi Lowongan Kerja pada Sistem Bursa Kerja Khusus Menggunakan Metode SAW Di SMK Nanda Riski Septania; Hakkun Elmunsyah; Utomo Pujianto
Edcomtech: Jurnal Kajian Teknologi Pendidikan Vol. 4 No. 2 (2019)
Publisher : Universitas Negeri Malang in collaboration with APSTPI and IPTPI

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.17977/um039v4i22019p120

Abstract

This study aims to produce recommendations in the form of a list of job vacancies in accordance with the expertise of students. The model used in research and development of student expertise mapping recommendations is the waterfall. The testing technique used is black box testing, which is used to test the system's functionality. The product was validated by two experts, namely Expert I and Expert II. The results of the validation carried out by expert I and expert II each received a 100% eligibility percentage. Then from the product trial results obtained a percentage of eligibility 97%. These results indicate that the level of product eligibility is valid because it exceeds the minimum limit of 76% for valid criteria, so the product can be said to be feasible and can be used.
Comparison of Naïve Bayes Algorithm and Decision Tree C4.5for Hospital Readmission Diabetes Patientsusing HbA1c Measurement Pujianto, Utomo; Setiawan, Asa Luki; Ar Rosyid, Harits; Salah, Ali M. Mohammad
Knowledge Engineering and Data Science
Publisher : citeus

Show Abstract | Download Original | Original Source | Check in Google Scholar

Abstract

Diabetes is a metabolic disorder disease in which the pancreas does not produce enough insulin or the body cannot use insulin produced effectively. The HbA1c examination, which measures the average glucose level of patients during the last 2-3 months, has become an important step to determine the condition of diabetic patients. Knowledge of the patient's condition can help medical staff to predict the possibility of patient readmissions, namely the occurrence of a patient requiring hospitalization services back at the hospital. The ability to predict patient readmissions will ultimately help the hospital to calculate and manage the quality of patient care. This study compares the performance of the Naïve Bayes method and C4.5 Decision Tree in predicting readmissions of diabetic patients, especially patients who have undergone HbA1c examination. As part of this study we also compare the performance of the classification model from a number of scenarios involving a combination of preprocessing methods, namely Synthetic Minority Over-Sampling Technique (SMOTE) and Wrapper feature selection method, with both classification techniques. The scenario of C4.5 method combined with SMOTE and feature selection method produces the best performance in classifying readmissions of diabetic patients with an accuracy value of 82.74 %, precision value of 87.1 %, and recall value of 82.7 %.
The Effect of Resampling on Classifier Performance: anEmpirical Study Pujianto, Utomo; Akbar, Muhammad Iqbal; Lassela, Niendhitta Tamia; Sutaji, Deni
Knowledge Engineering and Data Science
Publisher : citeus

Show Abstract | Download Original | Original Source | Check in Google Scholar

Abstract

An imbalanced class on a dataset is a common classification problem. The effect of using imbalanced class datasets can cause a decrease in the performance of the classifier. Resampling is one of the solutions to this problem. This study used 100 datasets from 3 websites: UCI Machine Learning, Kaggle, and OpenML. Each dataset will go through 3 processing stages: the resampling process, the classification process, and the significance testing process between performance evaluation values of the combination of classifier and the resampling using paired t-test. The resampling used in the process is Random Undersampling, Random Oversampling, and SMOTE. The classifier used in the classification process is Naïve Bayes Classifier, Decision Tree, and Neural Network. The classification results in accuracy, precision, recall, and fmeasure values are tested using paired t-tests to determine the significance of the classifier's performance from datasets that were not resampled and those that had applied the resampling. The paired t-test is also used to find a combination between the classifier and the resampling that gives significant results. This study obtained two results. The first result is that resampling on imbalanced class datasets can substantially affect the classifier's performance more than the classifier's performance from datasets that are not applied the resampling technique. The second result is that combining the Neural Network Algorithm without the resampling provides significance based on the accuracy value. Combining the Neural Network Algorithm with the SMOTE technique provides significant performance based on the amount of precision, recall, and f-measure. T
Co-Authors Abda Abda Achmad Murdiono Aguwin Ardi Pranata Aji Prasetya Wibawa Ali M. Mohammad Salah Angeline, Grace Anik Nur Handayani Anis Nikmatul Kasanah Annas Gading Pertiwi Annisa Putri Ayudhitama Aris Maulana Arya Yudhi Wijaya Asa Luki Setiawan Asmoro, Achmad Shoddiq Bayu Ayudhitama, Annisa Putri Bahalwan, Lugas Anegah Baihaqi, Dimas Imam Baihaqi, Dimas Imam Bintang Romadhon Binti Afifah Daniel Oranova Siahaan Deni Sutaji Dewandra, Aderyan Reynaldi Fahrezza Dewinta Nilan Sari Didik Dwi Prasetya Didin Rosyadi Dwi Jaelani, Mardian Elsa Dwi Rochmah Rachmanto Fauzan Cahya Arifin Filby, Brilliant Gradiyanto Radityo Kusumo Hakkun Elmunsyah Hammad, Jehad A. H. Harits Ar Rosyid Harits Ar Roysid Haviluddin Haviluddin Hery Widijanto Icca Astrina Izdihar, Zahra Nabila Joumil Aidil Saifuddin Khairunnisa, Raissa Araminta Kusumo, Gradiyanto Radityo Lassela, Niendhitta Tamia Lazuardi Noorca Rachmadi M. Zainal Arifin Martha, Jefry Aulia Meiga Ayu Ariyanti Muhamad Jauharul Fuady Muhammad Aqshal Muhammad Iqbal Akbar Muis Muhtadi Muladi Murad, Safwan Marwin Abdul Nafalski, Andrew Najahah, Vina Nanda Riski Septania Nanda Riski Septania Niendhitta Tamia Lassela Nur Hidayat, Wahyu Nurroby Wahyu Saputra Prananda Anugrah Putra Utama, Agung Bella Putri yuni Ristanti Raffi Taufik Gushardana Salah, Ali M. Mohammad Santoso, Priyo Aji Saputra, Anggie Wahyu Saputra, Irzan Tri Setiadi Cahyono Putro Setiawan, Asa Luki Swasono Raharjo, Swasono Triyanna Widiyaningtyas Utama, Agung Bella Putra Wahyu Sakti Gunawan Irianto Yudhistira, Moch Rajendra