p-Index From 2021 - 2026
8.753
P-Index
This Author published in this journals
All Journal IJCCS (Indonesian Journal of Computing and Cybernetics Systems) Jurnal Pseudocode Journal of Information Systems Engineering and Business Intelligence Sistemasi: Jurnal Sistem Informasi JOIV : International Journal on Informatics Visualization Sinkron : Jurnal dan Penelitian Teknik Informatika Jurnal RESTI (Rekayasa Sistem dan Teknologi Informasi) JURNAL MEDIA INFORMATIKA BUDIDARMA Jurnal Eksplora Informatika JITK (Jurnal Ilmu Pengetahuan dan Komputer) Techno Nusa Mandiri : Journal of Computing and Information Technology JOURNAL OF APPLIED INFORMATICS AND COMPUTING Jurnal Teknoinfo Jurnal Sisfokom (Sistem Informasi dan Komputer) Jurnal Infomedia JURNAL PengaMAS MATRIK : Jurnal Manajemen, Teknik Informatika, dan Rekayasa Komputer JURIKOM (Jurnal Riset Komputer) Jutisi: Jurnal Ilmiah Teknik Informatika dan Sistem Informasi Jurnal Informatika dan Rekayasa Elektronik Jurnal Riset Sistem Informasi dan Teknologi Informasi (JURSISTEKNI) Information System Journal (INFOS) JTECS : Jurnal Sistem Telekomunikasi Elektronika Sistem Kontrol Power Sistem dan Komputer Jurnal Pengabdian Mitra Masyarakat (JPMM) Malcom: Indonesian Journal of Machine Learning and Computer Science International Journal of Advanced Science Computing and Engineering SmartComp Fahma : Jurnal Informatika Komputer, Bisnis dan Manajemen Scientific Journal of Informatics SWAGATI: Journal of Community Service Edu Komputika Journal Jurnal Pengabdian Masyarakat Inovasi Indonesia
Claim Missing Document
Check
Articles

The Effect of Class Imbalance Handling on Datasets Toward Classification Algorithm Performance Cherfly Kaope; Yoga Pristyanto
MATRIK : Jurnal Manajemen, Teknik Informatika dan Rekayasa Komputer Vol. 22 No. 2 (2023)
Publisher : Universitas Bumigora

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.30812/matrik.v22i2.2515

Abstract

Class imbalance is a condition where the amount of data in the minority class is smaller than that of the majority class. The impact of the class imbalance in the dataset is the occurrence of minority class misclassification, so it can affect classification performance. Various approaches have been taken to deal with the problem of class imbalances such as the data level approach, algorithmic level approach, and cost-sensitive learning. At the data level, one of the methods used is to apply the sampling method. In this study, the ADASYN, SMOTE, and SMOTE-ENN sampling methods were used to deal with the problem of class imbalance combined with the AdaBoost, K-Nearest Neighbor, and Random Forest classification algorithms. The purpose of this study was to determine the effect of handling class imbalances on the dataset on classification performance. The tests were carried out on five datasets and based on the results of the classification the integration of the ADASYN and Random Forest methods gave better results compared to other model schemes. The criteria used to evaluate include accuracy, precision, true positive rate, true negative rate, and g-mean score. The results of the classification of the integration of the ADASYN and Random Forest methods gave 5% to 10% better than other models.
Investigating The Effectiveness of Various Convolutional Neural Network Model Architectures for Skin Cancer Melanoma Classification Rizky Hafizh Jatmiko; Yoga Pristyanto
MATRIK : Jurnal Manajemen, Teknik Informatika dan Rekayasa Komputer Vol. 23 No. 1 (2023)
Publisher : Universitas Bumigora

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.30812/matrik.v23i1.3185

Abstract

Melanoma is one of the most dangerous types of skin cancer. Since 2018, the number of skin cancer cases in the US has increased and exceeded 100,000. Melanoma is the third most common cancer in Indonesia, following womb cancer and breast cancer. Standard detection of melanoma skin cancer biopsy is costly and time-consuming. The purpose of this research is to build and compare melanoma skin cancer detection using various Convolutional Neural Network method. This research used four CNN model architectures methods, VGG-16, LeNet, Xception, and MobileNet. The dataset for this research is image data that consists of 9605 data divided into benign and malignant classes. The data will be augmented to increase its quantity. After that, the data will be trained using four CNN architecture models and evaluated using the confusion matrix. The result of this study is that Xception model has the best accuracy and the lowest loss, with 93% accuracy and 19% loss, with precision 93%, recall 93,5%, and f1-score 93%. Whereas the other model, VGG-16 gives 90 % accuracy, 27% loss, LeNet 89,7% accuracy, 28% loss, and mobileNet 90,8% accuracy and 22,5% loss.
Stock Price Prediction Using SVR: A Feature Engineering and Hyperparameter Tuning Approach Alfian Ramadhan; Yoga Pristyanto; Anggit Dwi Hartanto; Donni Prabowo
Jurnal RESTI (Rekayasa Sistem dan Teknologi Informasi) Vol 10 No 3 (2026): June 2026
Publisher : Ikatan Ahli Informatika Indonesia (IAII)

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.29207/resti.v10i3.7249

Abstract

Stock price prediction in Indonesia's volatile mining sector poses significant forecasting challenges driven by commodity price dynamics and structural market shifts. This study proposes a systematic prediction framework for PT Indo Tambangraya Megah Tbk (ITMG.JK) integrating technical and market-derived non-technical feature engineering, LightGBM-based feature selection, multilevel TimeSeriesSplit cross-validation, and hyperparameter optimization. Support Vector Regression (SVR) is benchmarked against LightGBM, XGBoost, and Random Forest under 5-fold, 10-fold, and 15-fold schemes. SVR achieves the best performance at 10-fold, with RMSE of 0.0121, MAE of 0.0090, MAPE of 1.1457%, and R² of 0.9249. Generalization experiments across four additional stocks in banking, automotive, and mining sectors confirm SVR's robustness, maintaining R² above 0.89 and MAPE below 2.65% in all cases while tree-based models produce negative R² on certain datasets. Statistical validation via Wilcoxon signed-rank test (p < 0.05) and Cohen's d (|d| > 0.8) confirms the significance of SVR's advantage. These findings indicate that SVR consistently outperforms the evaluated models under the proposed experimental framework.
Evaluasi Komparatif Algoritma Machine Learning dalam Analisis Sentimen Program Makan Bergizi Gratis di Media Sosial X Angelina Putri Ariani; Norhikmah; Yoga Pristyanto
JURIKOM (Jurnal Riset Komputer) Vol. 13 No. 3 (2026): Juni 2026
Publisher : Universitas Budi Darma

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.30865/jurikom.v13i3.9741

Abstract

The analysis of public opinion on social media platforms through sentiment analysis plays a crucial role in understanding how the public responds to government policies, including the Free Nutritious Meal Program (MBG). However, the imbalanced nature of social media data and the use of informal language such as sarcasm pose challenges in the sentiment classification process. Therefore, this study aims to examine public perceptions of the MBG program on platform X while also evaluating the effectiveness of several machine learning algorithms in categorizing sentiment. The dataset used in this study consists of 2,000 comments collected between February and April 2026. The data were labeled using a lexicon-based approach and processed through preprocessing and feature extraction using TF-IDF. The classification process was carried out using six algorithms: Naïve Bayes, K-Nearest Neighbor (K-NN), Support Vector Machine (SVM), Decision Tree, Random Forest, and Logistic Regression. The results show that Random Forest achieved the highest accuracy, reaching 92%, supported by a cross-validation score of 89%, indicating strong model stability. Based on the classification results, public sentiment is predominantly neutral at 66.3%, followed by negative sentiment at 22.6% and positive sentiment at 11.1%. These findings suggest that public opinion toward the MBG program tends to be neutral, with a stronger inclination toward criticism than support. Furthermore, the results highlight the importance of selecting appropriate algorithms to improve the accuracy of sentiment analysis on complex and imbalanced textual data
Prediksi Harga Rumah Menggunakan XGBoost Berbasis Optuna Hyperparameter Optimization: House Price Prediction using XGBoost Based on Optuna Hyperparameter Optimization Natasaskara, Nandana Ayudya; Pristyanto, Yoga; Fajri, Ika Nur
MALCOM: Indonesian Journal of Machine Learning and Computer Science Vol. 6 No. 3 (2026): MALCOM July 2026
Publisher : Institut Riset dan Publikasi Indonesia

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.57152/malcom.v6i3.2704

Abstract

Fluktuasi dan kompleksitas atribut pasar properti membuat prediksi harga rumah sulit dan kurang akurat jika hanya mengandalkan parameter algoritma bawaan. Kinerja optimal algoritma Machine learning seperti Extreme Gradient Boosting (XGBoost) sangat bergantung pada pengaturan hyperparameter yang tepat, namun banyak penelitian sebelumnya mengabaikan optimasi atau menggunakan metode konvensional yang tidak efisien. Untuk mengatasi masalah tersebut, penelitian ini mengusulkan model XGBoost yang dioptimasi secara dinamis menggunakan kerangka kerja Optuna Hyperparameter Optimization. Optuna, yang bekerja berdasarkan optimasi Bayesian, secara cerdas dan efisien mengeksplorasi ruang parameter guna menemukan konvergensi yang lebih cepat. Hasil eksperimen membuktikan bahwa integrasi Optuna berhasil meningkatkan keakurasian prediksi secara signifikan. Model XGBoost berbasis Optuna menghasilkan performa yang lebih unggul dengan peningkatan skor R² dari 0.8204 menjadi 0.8263, serta berhasil menekan tingkat kesalahan di mana RMSE turun menjadi Rp 301.132.090,80, MAE menjadi Rp 196.869.100,79, dan MAPE menyusut menjadi 16,53%. Pendekatan ini terbukti lebih tangguh, stabil, dan presisi dibandingkan model tanpa optimasi (baseline) dalam memetakan pola harga yang non-linear. Meskipun akurasi meningkat, hal ini menuntut waktu komputasi pelatihan yang jauh lebih tinggi, yakni melonjak drastis menjadi 3.550,34 detik. Dataset akhir yang digunakan dalam penelitian ini berjumlah 16.674 catatan, yang diperoleh setelah proses preprocessing dan eliminasi outlier secara ekstensif dari 40.200 catatan awal.
Predictive Modeling of Urban Air Quality Using Machine Learning Dahlan, Akhmad; Fatta, Hanif Al; Farida, Lilis Dwi; Utama, Hastari; Pristyanto, Yoga
IJCCS (Indonesian Journal of Computing and Cybernetics Systems) Vol 20, No 3 (2026): July
Publisher : IndoCEISS in colaboration with Universitas Gadjah Mada, Indonesia.

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.22146/ijccs.117546

Abstract

Air pollution is a critical environmental concern adversely affecting public health in urban areas worldwide. Accurate prediction of air quality enables timely intervention by environmental agencies and policymakers. This study presents a comprehensive comparative analysis of ensemble machine learning methods for predicting urban air quality using the publicly available UCI Air Quality Dataset, containing 9,358 hourly instances of chemical sensor measurements collected in a heavily polluted Italian city from March 2024 to February 2025. Five algorithms were evaluated: Random Forest (RF), Gradient Boosting Machine (GBM), XGBoost, LightGBM, and CatBoost, against Ridge Regression and SVR baselines. A systematic pipeline encompassing preprocessing, temporal feature engineering, Bayesian hyperparameter optimization via Optuna, and time-series cross-validation was implemented. Results demonstrate XGBoost achieved the best performance with RMSE = 2.14, MAE = 1.58, and R² = 0.9312. SHAP-based feature importance analysis revealed CO(GT) lag features, C6H6(GT), and NOx(GT) as the most influential predictors.
Performance Evaluation of Machine Learning Models for Soil Fertility Classification Based on the Indian Soil Fertility Dataset Yoga Pristyanto; Ibrahim Aji Fajar Romadhon; Anggit Ferdita Nugraha; Atik Nurmasani; Irma Rofni Wulandari
Edu Komputika Journal Vol. 12 No. 1 (2025): Edu Komputika Journal
Publisher : Universitas Negeri Semarang

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.15294/edukom.v12i1.10317

Abstract

Rice farming productivity worldwide has been declining due to improper soil management practices, including excessive chemical fertilizer use and irregular irrigation. The main challenge lies in accurately classifying soil fertility levels to support optimal land use and reduce resource waste, especially when dealing with imbalanced datasets. This study aims to compare the performance of single classifiers and ensemble classifiers in classifying soil fertility. The single classifiers used include K-Nearest Neighbor (KNN), Naive Bayes, Decision Tree, Support Vector Machine (SVM), and Artificial Neural Network (ANN), while the ensemble classifiers include Random Forest and XGBoost. The Indian Soil Fertility Dataset, obtained from Kaggle, contains 880 samples with 12 features and 1 output class. The research methodology involved data acquisition, preprocessing, data splitting, standardization, and classification, with performance evaluation conducted using a confusion matrix. The results show that ensemble classifiers, particularly Random Forest and XGBoost, outperform single classifiers in imbalanced datasets, achieving accuracy, precision, recall, and F1-score values exceeding 92%-95% across all split scenarios. The findings conclude that Random Forest and XGBoost can serve as reliable models for assisting farmers and agricultural experts in evaluating soil conditions, minimizing unnecessary fertilizer usage, and improving rice farming productivity globally.
A Hybrid Intersection Filtering and Recursive Feature Elimination Technique for Efficient Feature Reduction in High Dimensional Datasets Akhmad Dahlan; Yoga Pristyanto; Anggit Ferdita Nugraha; Rifda Faticha Alfa Aziza; Ibnu Hadi Purwanto
Jurnal RESTI (Rekayasa Sistem dan Teknologi Informasi) Vol 10 No 2 (2026): April 2026
Publisher : Ikatan Ahli Informatika Indonesia (IAII)

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.29207/resti.v10i2.7396

Abstract

High-dimensional datasets are commonly encountered in real-world machine learning applications and often degrade classification performance due to redundant and irrelevant features. In addition, the presence of excessive features increases computational complexity and processing time. Feature selection is therefore a crucial preprocessing step to improve model accuracy and efficiency. This study proposes a hybrid feature selection approach called Intersection Filtering based on Recursive Feature Elimination with Cross-Validation (IF-RFECV), which integrates wrapper-based and filter-based strategies to obtain a stable and optimal subset of features. The proposed method first applies Recursive Feature Elimination with Cross-Validation (RFECV) using multiple classification models to rank and select relevant features. Subsequently, an intersection filtering mechanism is employed to identify features that are consistently selected across different RFECV-based models, thereby reducing model-dependent bias and improving feature robustness. The effectiveness of IF-RFECV is evaluated using four benchmark datasets with varying dimensionality obtained from the KEEL and UCI repositories. Several classification algorithms, including Gradient Boosting, K-Nearest Neighbor, Naïve Bayes, Decision Tree, Random Forest, and Support Vector Machine, are used to assess model performance. Experimental results demonstrate that IF-RFECV produces a more compact feature subset compared to conventional RFECV while achieving superior performance in terms of accuracy, precision, recall, and F1-score on most datasets, particularly those with higher dimensionality. Although IF-RFECV requires slightly higher computational time due to its two-stage process, the performance gains and improved generalization justify this trade-off. These findings indicate that IF-RFECV is an effective and robust feature selection technique for high-dimensional classification problems.
Cross-Dataset Evaluation of Boosting Models for Hypertension Prediction Bety Wulan Sari; Dewi Ayu Murtiningsih; Donni Prabowo; Yoga Pristyanto; Ika Nur Fajri; Ike Verawati
Jurnal RESTI (Rekayasa Sistem dan Teknologi Informasi) Vol 10 No 4 (2026): August 2026 (in progress)
Publisher : Ikatan Ahli Informatika Indonesia (IAII)

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.29207/resti.v10i4.7620

Abstract

Hypertension remains a significant global risk factor for cardiovascular disease and related mortality, necessitating reliable early risk prediction models. Although boosting algorithms have demonstrated strong performance in structured medical data, limited studies have examined their consistency across heterogeneous datasets. This study aims to evaluate the cross-dataset performance and stability of three boosting models, such as XGBoost, LightGBM, and CatBoost, for hypertension prediction under multiple train–test split ratios. Two independent structured datasets were analyzed using 60:40, 70:30, 80:20, and 90:10 splits. To identify the optimal hyperparameters, grid search was performed using repeated stratified 5-fold cross-validation with three repetitions. Model effectiveness was measured using the evaluation metrics of accuracy, precision, recall, F1-score, and AUC. Results show that Dataset 1 gained consistently high predictive performance (accuracy > 0.98; AUC ≈ 1.00), indicating strong and well-separated predictive signals, whereas Dataset 2 demonstrated substantially lower discriminative ability (accuracy ≈ 0.71–0.72; AUC ≈ 0.50), suggesting limited predictive structure. Across both datasets, CatBoost consistently obtained the highest accuracy, particularly at the 90:10 split ratio. These findings demonstrate that dataset characteristics critically determine model effectiveness and that among the evaluated boosting algorithms, CatBoost delivered the strongest overall predictive performance.
Aceh Province Tourism Destination Recommendation System using Content based Filtering Method Ariefhan Maulana; Arif Nur Rohman; Yoga Pristyanto
SISTEMASI Vol 15, No 2 (2026): Sistemasi: Jurnal Sistem Informasi
Publisher : Universitas Islam Indragiri

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.32520/stmsi.v15i2.5500

Abstract

Tourists often experience difficulties in finding tourist destinations in Aceh Province that match their content preferences and are geographically close to their location. This study aims to develop a tourism destination recommendation system in Aceh Province using a Content-Based Filtering approach with the Cosine Similarity algorithm and the Haversine Formula. The dataset consists of 119 tourist destinations, including attributes such as destination name, destination description, and geographical coordinates (latitude and longitude). The research process began with text data preprocessing, which included case folding, punctuation removal, tokenization, duplicate word removal, stopword removal, and stemming. Next, the similarity between destinations was calculated using the Cosine Similarity algorithm based on tourism content descriptions, while the Haversine Formula was applied to measure the geographical distance between the user’s location and the tourist destinations. The results indicate that the developed system is able to provide relevant tourism destination recommendations by simultaneously considering content relevance and geographical proximity. Therefore, the system can assist tourists in selecting destinations that best match their preferences.
Co-Authors Acihmah Sidauruk Aditya Yoga Pratama Afrig Aminuddin Aisha Shakila Iedwan Akhmad Dahlan Alfian Ramadhan Alvin Rahman Al Musyaffa Andi Sunyoto Angelina Putri Ariani Anggi Thoat Ariyanto Anggit Dwi Hartanto Anggit Dwi Hartanto Anggit Dwi Hartanto Anggit Dwi Hartanto, Anggit Dwi Anggita, Sharazita Dyah Anna Baita Ariefhan Maulana arif nur rohman Arif Nur Rohman Arif Nur Rohman Asti Astuti, Ika Atik Nurmasani ATIK NURMASANI Atik Nurmasani Barus, Herianta Bety Wulan Sari Bety Wulan Sari, Bety Wulan Bligania Bligania Cherfly Kaope Dewi Ayu Murtiningsih Donni Prabowo Donni Prabowo, Donni Dwi Hartanto, Anggit Dyah Anggita, Sharazita Eli Pujastuti, Eli Eza Nanda Fadhilah Dwi Ananda Fajri, Ika Nur Fauzy, Marwan Noor Gagah Gumelar Gita Cahyani Hanif Al Fatta Heri Sismoro Hidayat, Kardilah Rohmat Ibnu Hadi Purwanto Ibnu Hadi Purwanto Ibrahim Aji Fajar Romadhon Iedwan, Aisha Shakila Ike Verawati Ikmah Ikmah Irfan Pratama Irma Rofni Wulandari Istikomah Khoiruddin, Lukman Kono, Maria Fatima Kristianti, Fanny Novatriana LILIS DWI FARIDA Lucky Adhikrisna Wirasakti Mambaul Hisam Marcheilla Trecya Anindita Mauliza, Nia Mukarabiman, Zulfikar Mulia Sulistiyono Natasaskara, Nandana Ayudya Nia Mauliza Nia Mauliza Norhikmah Nugraha, Anggit Ferdita Nuri Cahyono Nurindah A Amari Nurwijayanti Purwati, Sintia Eka Putra, Frahma Aditya Rahman Saputra, Rahman Rifda Faticha Alfa Aziza Rizky Hafizh Jatmiko Rohmad Fajarudin Rohman, Arif Nur Romadhon, Ibrahim Aji Fajar Rospita, Andri Sabella, Cindy Dinda Sifa’ul Husna, Siti Okta Sumarni Adi Utama, Hastari Windarni, Vikky Aprelia Wirantanu, Dipa Wirasakti, Lucky Adhikrisna Wiwi Widayani Yanuar Nur Kholik Yogata Rama Guninta Yudiyanto, Muhammad Resa Arif Yuli Astuti Zein, Aditya Ahmad