Claim Missing Document
Check
Articles

Found 31 Documents
Search

An Empirical Study of Cross-Project and Within-Project Performance in Software Defect Prediction Models Using Tree-Based and Boosting Classifiers Raidra Zeniananto; Herteno, Rudy; Radityo Adi Nugroho; Andi Farmadi; Setyo Wahyu Saputro
Indonesian Journal of Electronics, Electromedical Engineering, and Medical Informatics Vol. 7 No. 3 (2025): August
Publisher : Jurusan Teknik Elektromedik, Politeknik Kesehatan Kemenkes Surabaya, Indonesia

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.35882/ijeeemi.v7i3.95

Abstract

Software Defect Prediction (SDP) is a vital process in modern software engineering aimed at identifying faulty components in the early stages of development. In this study, we conducted a comprehensive evaluation of two widely employed SDP approaches, Within-Project Software Defect Prediction (WP-SDP) and Cross-Project Software Defect Prediction (CP-SDP), using identical preprocessing steps to ensure an objective comparison. We utilized the NASA MDP dataset, where each project was split into 70% training and 30% testing data, and applied three distinct resampling strategies—no sampling, oversampling, and undersampling—to address the challenge of class imbalance. Five classification algorithms were examined, including Support Vector Machine (SVM), Random Forest (RF), Gradient Boosting (GB), XGBoost (XGB), and LightGBM (LGBM). Performance was measured primarily using Accuracy and Area Under the Curve (AUC) metrics, resulting in 360 experimental outcomes. Our findings revealed that WP-SDP, combined with oversampling and Random Forest, demonstrated superior predictive capability on most projects, achieving an Accuracy of 89.92% and an AUC of 0.931 on PC4. Nonetheless, CP-SDP excelled in certain small-scale projects (e.g., MW1), underscoring its potential when local historical data is scarce but inter-project characteristics remain sufficiently similar. This study’s results underscore the importance of selecting a prediction scheme tailored to specific project attributes, class imbalance levels, and available historical data. By establishing a standardized methodological framework, our work contributes to a clearer understanding of the strengths and limitations of WP-SDP and CP-SDP, paving the way for more effective defect detection strategies and improved software quality.
Implementation of Ant Colony Optimization in Obesity Level Classification Using Random Forest Wardana, Muhammad Difha; Budiman, Irwan; Indriani, Fatma; Nugrahadi, Dodon Turianto; Saputro, Setyo Wahyu; Rozaq, Hasri Akbar Awal; Yıldız, Oktay
Jurnal Teknik Informatika (Jutif) Vol. 6 No. 5 (2025): JUTIF Volume 6, Number 5, Oktober 2025
Publisher : Informatika, Universitas Jenderal Soedirman

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.52436/1.jutif.2025.6.5.4696

Abstract

Obesity is a pressing global health issue characterized by excessive body fat accumulation and associated risks of chronic diseases. This study investigates the integration of Ant Colony Optimization (ACO) for feature selection in obesity-level classification using Random Forests. Results demonstrate that feature selection significantly improves classification accuracy, rising from 94.49% to 96.17% when using ten features selected by ACO. Despite limitations, such as challenges in tuning parameters like alpha (α), beta (β), and evaporation rate in ACO techniques, the study provides valuable insights into developing a more efficient obesity classification system. The proposed approach outperforms other algorithms, including KNN (78.98%), CNN (82.00%), Decision Tree (94.00%), and MLP (95.06%), emphasizing the importance of feature selection methods like ACO in enhancing model performance. This research addresses a critical gap in intelligent healthcare systems by providing the first comprehensive study of ACO-based feature selection specifically for obesity classification, contributing significantly to medical informatics and computer science. The findings have immediate practical implications for developing automated diagnostic tools that can assist healthcare professionals in early obesity detection and intervention, potentially reducing healthcare costs through improved diagnostic efficiency and supporting digital health transformation in clinical settings. Furthermore, the study highlights the broader applicability of ACO in various classification tasks, suggesting that similar techniques could be used to address other complex health issues, ultimately improving diagnostic accuracy and patient outcomes.
Accurate Skin Tone Classification for Foundation Shade Matching using GLCM Features-K-Nearest Neighbor Algorithm Syahputra, Muhammad Reza; Mazdadi, Muhammad Itqan; Budiman, Irwan; Farmadi, Andi; Saputro, Setyo Wahyu; Rozaq, Hasri Akbar Awal; Sutaji, Deni
Jurnal Teknik Informatika (Jutif) Vol. 6 No. 5 (2025): JUTIF Volume 6, Number 5, Oktober 2025
Publisher : Informatika, Universitas Jenderal Soedirman

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.52436/1.jutif.2025.6.5.4723

Abstract

Foundation shade matching remains a significant challenge in the beauty industry, particularly in Indonesia where consumers exhibit three distinct skin tone categories: ivory white, amber yellow, and tan. Manual foundation selection often results in mismatched shades, leading to customer dissatisfaction. This study presents a novel automated skin tone classification system combining Gray Level Co-Occurrence Matrix (GLCM) feature extraction with the K-Nearest Neighbor (KNN) algorithm. The GLCM method extracts four key texture features (contrast, homogeneity, energy, and entropy) from facial images, while KNN performs classification. A comprehensive dataset of 963 facial images was used, with 770 training and 193 test samples collected under controlled lighting conditions. After testing K values from 1 to 15, the optimal K=1 achieved 75.65% accuracy. Compared to baseline color histogram methods (60% accuracy), our GLCM-KNN approach demonstrates 15.65% improvement in classification performance. This research contributes to computer vision applications in beauty technology, enabling the development of mobile applications for virtual foundation try-on and personalized product recommendations. The findings have significant implications for the cosmetics industry, particularly for automated cosmetic shade matching systems and enhanced customer experience in online beauty retail. Further research is recommended to explore deep learning approaches and expand dataset diversity to improve accuracy.
Evaluation of User Experience in the Banjarbaru Disdukcapil Public Service Application Using User Experience Questionnaire and System Usability Scale Martalisa, Asri; Wahyu Saputro, Setyo; Turianto Nugrahadi, Dodon; Abadi, Friska; Budiman, Irwan
Scientific Journal of Informatics Vol. 11 No. 4: November 2024
Publisher : Universitas Negeri Semarang

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.15294/sji.v11i4.13780

Abstract

Purpose: Dukcapil Banjarbaru is an online-based government agency application used for various public services. According to the complaint report from Disdukcapil Banjarbaru, several users have reported similar problems and difficulties. The application has received a rating of 3.3 stars from approximately 24.000 users on the Google Play Store. Therefore, researchers conducted a user experience analysis using the UEQ methods and a usability evaluation using the SUS methods. Methods: This research analyzes user experience in applications using the UEQ to identify issues faced by users and evaluate usability through the System Usability Scale. The UEQ method is chosen for its efficiency and simplicity in assessing user experience within an application. The SUS method is employed because it is an effective approach for obtaining reliable statistical data and generating clear and accurate scores. Result: The UEQ benchmark results show that the scales for Attractiveness (1.59), Efficiency (1.68), Accuracy (1.66), and Stimulation (1.54) are categorized as "Good." The scales for clarity (1.37) and novelty (0.80) are classified as "Above Average." Meanwhile, the SUS score of 65 positions the application within the "acceptable" category for the acceptability range, the "D" category on the grade scale, and the "OK" category for adjective ratings. This indicates that while the Banjarbaru Dukcapil application has good usability, it requires improvements based on the total SUS score, which reveals several critical areas with scores below the average (258.4). Novelty: In this research, solutions for improvements are provided to Disdukcapil based on each aspect to improve the quality of the application, thereby offering better services to users.
Dimensionality Reduction Using Principal Component Analysis and Feature Selection Using Genetic Algorithm with Support Vector Machine for Microarray Data Classification Kartini, Dwi; Badali, Rahmat Amin; Muliadi, Muliadi; Nugrahadi, Dodon Turianto; Indriani, Fatma; Saputro, Setyo Wahyu
Indonesian Journal of Electronics, Electromedical Engineering, and Medical Informatics Vol. 7 No. 1 (2025): February
Publisher : Jurusan Teknik Elektromedik, Politeknik Kesehatan Kemenkes Surabaya, Indonesia

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.35882/mr7x9713

Abstract

DNA microarray is used to analyze gene expression on a large scale simultaneously and plays a critical role in cancer detection. The creation of a DNA microarray starts with RNA isolation from the sample, which is then converted into cDNA and scanned to generate gene expression data. However, the data generated through this process is highly dimensional, which can affect the performance of predictive models for cancer detection. Therefore, dimensionality reduction is required to reduce data complexity. This study aims to analyze the impact of applying Principal Component Analysis (PCA) for dimensionality reduction, Genetic Algorithm (GA) for feature selection, and their combination on microarray data classification using Support Vector Machine (SVM). The datasets used are microarray datasets, including breast cancer, ovarian cancer, and leukemia. The research methodology involves preprocessing, PCA for dimensionality reduction, GA for feature selection, data splitting, SVM classification, and evaluation. Based on the results, the application of PCA dimensionality reduction combined with GA feature selection and SVM classification achieved the best performance compared to other classifications. For the breast cancer dataset, the highest accuracy was 73.33%, recall 0.74, precision 0.75, and F1 score 0.73. For the ovarian cancer dataset, the highest accuracy was 98.68%, recall 0.98, precision 0.99, and F1 score 0.99. For the leukemia dataset, the highest accuracy was 95.45%, recall 0.94, precision 0.97, and F1 score 0.95. It can be concluded that combining PCA for dimensionality reduction with GA for feature selection in microarray classification can simplify the data and improve the accuracy of the SVM classification model. The implications of this study emphasize the effectiveness of applying PCA and GA methods in enhancing the classification performance of microarray data.
Hybrid Feature Selection and Balancing Data Approach for Improved Software Defect Prediction Febrian, Muhamad Michael; Saputro, Setyo Wahyu; Saragih, Triando Hamonangan; Abadi, Friska; Herteno, Rudy
Indonesian Journal of Electronics, Electromedical Engineering, and Medical Informatics Vol. 7 No. 2 (2025): May
Publisher : Jurusan Teknik Elektromedik, Politeknik Kesehatan Kemenkes Surabaya, Indonesia

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.35882/ijeeemi.v7i2.67

Abstract

Software Defect Prediction (SDP) plays a vital role in identifying defects within software modules. Accurate early detection of software defects can reduce development costs and enhance software reliability. However, SDP remains a significant challenge in the software development lifecycle. This study employs Particle Swarm Optimization (PSO) and addresses several challenges associated with its application, including noisy attributes, high-dimensional data, and imbalanced class distribution. To address these challenges, this study proposed a hybrid filter-based feature selection and class balancing method. The feature selection process incorporates Chi-Square (CS), Correlation-Based Feature Selection (CFS), and Correlation Matrix-Based Feature Selection (CMFS), which have been proven effective in reducing noisy and redundant attributes. Additionally, the Synthetic Minority Over-sampling Technique (SMOTE) is applied to mitigate class imbalance in the dataset. The K-Nearest Neighbors (KNN) algorithm is employed as the classification model due to its simplicity, non-parametric nature, and suitability for handling the feature subsets produced. Performance evaluation is conducted using the Area Under Curve (AUC) metric with a significance threshold of 0.05 to assess classification capability.  The proposed method achieved an AUC of 0.872, demonstrating its effectiveness in enhancing predictive performance. The proposed method was also superior to other combinations such as PSO SMOTE (0.0043), PSO SMOTE CS (0.0091), PSO SMOTE CFS (0.0111), and PSO SMOTE CFS CMFS (0.0007). The findings of this study show that the proposed method significantly enhances the efficiency and accuracy of PSO in software defect prediction tasks. This hybrid strategy demonstrates strong potential as a robust solution for future research and application in predictive software quality assurance.
Application of Adaboost Algorithm with SMOTE and Optuna Techniques in Sleep Disorder Classification Anshory, Muhammad Naufal; Mazdadi, Muhammad Itqan; Saragih, Triando Hamonangan; Budiman, Irwan; Saputro, Setyo Wahyu
Indonesian Journal of Electronics, Electromedical Engineering, and Medical Informatics Vol. 7 No. 2 (2025): May
Publisher : Jurusan Teknik Elektromedik, Politeknik Kesehatan Kemenkes Surabaya, Indonesia

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.35882/ijeeemi.v7i2.99

Abstract

Data imbalance is a serious challenge in developing machine learning models for sleep disorder classification. When models are trained on an uneven distribution of classes, classification performance for minority classes such as insomnia and sleep apnea is often low. As a result, the overall accuracy may seem elevated, yet the sensitivity to important cases to be weak. Therefore, this research aims to design and develop a robust sleep disorder classification model with the AdaBoost algorithm, with improved performance through the integration of two main approaches, namely data balancing technique utilizing SMOTE and hyperparameter optimization using Optuna. This research contributes by showing that the combination of the two approaches can significantly improve model performance, not only in terms of global accuracy, but also accuracy on previously overlooked minority classes. The dataset utilized is the Sleep Health and Lifestyle Dataset which consists of 374 synthesized data and is divided into three categories: insomnia, sleep apnea, and none. This method stages include data preprocessing, data division using train-test split (80:20), application of SMOTE to balance the class distribution, hyperparameter tuning using Optuna, and model training with the AdaBoost algorithm. Evaluation was performed using classification metrics: accuracy, precision, recall, and F1-score. Results showed that mix of SMOTE and Optuna yielded the best results, accuracy 90.6%, F1-score 0.83871 for insomnia, and 0.81250 for sleep apnea. This performance was consistently superior to scenarios with no SMOTE or no tuning. This confirms the importance of using combination strategies to obtain fair and accurate classification on medical data. Future research is recommended to use real datasets as well as test the capabilities of this research on other models such as XGBoost or LightGBM.
Multi-Criteria Decision Making dalam Seleksi Fitur Ensemble untuk Prediksi Cacat Perangkat Lunak Fikri, Muhammad; Herteno, Rudy; Adi Nugroho, Radityo; Wahyu Saputro, Setyo; Abadi, Friska
Jurnal Teknologi Informasi dan Ilmu Komputer Vol 12 No 6: Desember 2025
Publisher : Fakultas Ilmu Komputer, Universitas Brawijaya

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.25126/jtiik.2025125

Abstract

Prediksi cacat perangkat lunak merupakan upaya strategis dalam meningkatkan kualitas produk melalui identifikasi dini modul yang berpotensi cacat. Kinerja prediksi dipengaruhi oleh pemilihan fitur, karena informasi yang berlebihan dan tidak relevan dapat mempengaruhi kualitas pembelajaran model. Seleksi fitur ensemble dinilai efektif dalam menyeleksi fitur yang relevan dengan menggabungkan beberapa metode seleksi fitur berbasis filter. Diperlukan mekanisme integrasi untuk menyatukan hasil dari empat teknik filter—Mutual Information, Fisher Score, Uncertainty dan Relief. Penelitian ini membandingkan empat metode Multi‑Criteria Decision Making—TOPSIS, VIKOR, EDAS, dan WASPAS—yang bekerja dengan merangking nilai relevansi fitur hasil seleksi filter tersebut. Sepuluh fitur teratas dari tiap metode kemudian dievaluasi menggunakan model Random Forest dengan metrik AUC melalui K‑Fold cross‑validation. Dari 12 dataset NASA MDP yang diuji, TOPSIS menunjukkan kinerja paling konsisten dan terbaik dengan nilai rata-rata AUC sebesar 0,8038. Temuan ini menegaskan pentingnya pemilihan metode integrasi yang tepat dalam meningkatkan akurasi prediksi cacat perangkat lunak dan memberikan panduan bagi pengembangan model yang lebih efektif.   Abstract Software defect prediction is a strategic effort to improve product quality through early identification of potentially defective modules. Prediction performance is influenced by feature selection, because redundant and irrelevant information can affect the quality of model learning. Ensemble feature selection is considered effective in selecting relevant features by combining several filter-based feature selection methods. An integration mechanism is needed to unify the results of four filter techniques—Mutual Information, Fisher Score, Uncertainty and Relief. This study compares four Multi-Criteria Decision Making methods—TOPSIS, VIKOR, EDAS, and WASPAS—which work by ranking the relevance values ​​of the filter-selected features. The top ten features from each method are then evaluated using the Random Forest model with the AUC metric through K-Fold cross-validation. Of the 12 NASA MDP datasets tested, TOPSIS showed the most consistent and best performance with an average AUC value of 0.8038. These findings emphasize the importance of choosing the right integration method in improving the accuracy of software defect prediction and provide guidance for the development of more effective models.
Dynamic Decay Adjustment in Radial Basis Function Networks: Does It Improve Software Defect Prediction? Kamil, Hawariul; Faisal, Mohammad Reza; Farmadi, Andi; Hertono, Rudy; Saputro, Setyo Wahyu
International Journal of Electronics and Communications Systems Vol. 5 No. 2 (2025): International Journal of Electronics and Communications System
Publisher : Universitas Islam Negeri Raden Intan Lampung, Indonesia

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.24042/ijecs.v5i2.29288

Abstract

Software quality depends heavily on the early detection of potentially defective modules, yet the complexity of software metrics and class imbalance often leads to inconsistent prediction performance. This study aims to compare the effectiveness of Radial Basis Function Neural Network (RBFNN) and RBFNN with Dynamic Decay Adjustment (RBFNN-DDA) in predicting software defects using five NASA PROMISE datasets (CM1, KC1, MC1, MW1, and PC1). The research employed quantitative experimentation through data normalization, a 70 to 30 train–test split, and model evaluation across maximum iterations ranging from 200 to 1,000. Model performance was assessed using Accuracy, Precision, Recall, F1 Score, and AUC. The results indicate that RBFNN provides higher Recall and F1 Score, making it better at identifying defective modules, although its performance is less stable. Meanwhile, RBFNN-DDA yields more consistent performance with higher Precision, Accuracy, and AUC on imbalanced datasets, albeit with lower Recall. Both models reached performance saturation at 200 until 400 iterations, showing minimal improvement at higher iteration counts. The findings imply the need for balancing sensitivity and stability when selecting defect prediction models, particularly in environments with severe class imbalance
Analisis Sentimen Ulasan Media Sosial UMKM Kuliner dengan Pendekatan Lexicon-Based dan Kosakata Khusus Setyo Wahyu Saputro; Friska Abadi; Radityo Adi Nugroho
Jurnal Informatika Polinema Vol. 12 No. 2 (2026): Vol. 12 No. 2 (2026)
Publisher : UPT P2M State Polytechnic of Malang

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.33795/jip.v12i2.9302

Abstract

UMKM kuliner di Kalimantan Selatan memanfaatkan media sosial sebagai sarana utama untuk mengetahui opini pelanggan, namun jumlah komentar yang sangat besar menyulitkan pelaku usaha untuk menelaahnya secara manual. Kondisi ini menegaskan perlunya pendekatan analisis sentimen yang mampu mengolah data ulasan secara efisien serta sesuai dengan karakteristik bahasa lokal. Penelitian ini bertujuan mengembangkan metode analisis sentimen berbasis lexicon yang diperkaya dengan kosakata domain-spesifik kuliner dan bahasa Banjar agar hasil klasifikasi lebih akurat dan kontekstual. Data penelitian diperoleh dari 3.500 komentar publik di Instagram dan TikTok. Tahap preprocessing mencakup case folding, pembersihan karakter khusus, tokenisasi, stopword removal, normalisasi, dan stemming. Selanjutnya, InSet Lexicon disempurnakan melalui penyuntikan kosakata baru serta penyesuaian bobot kata sesuai konteks kuliner lokal. Hasil analisis menunjukkan distribusi sentimen terdiri dari 2.050 komentar positif (58,57%), 934 komentar netral (26,69%), dan 516 komentar negatif (14,74%). Evaluasi menunjukkan peningkatan akurasi signifikan setelah perluasan lexicon, yaitu 93,49% untuk sentimen negatif, 94,64% untuk netral, dan 96,94% untuk positif, dibandingkan akurasi awal yang berkisar antara 51–73%. Temuan ini membuktikan bahwa pengayaan lexicon menggunakan kosakata lokal dan domain-spesifik secara substansial meningkatkan performa analisis sentimen. Pendekatan ini memberikan solusi praktis dan terjangkau bagi UMKM untuk memahami opini pelanggan secara lebih representatif, serta dapat dimanfaatkan dalam pengambilan keputusan strategis dan perbaikan kualitas layanan maupun promosi produk kuliner.