Claim Missing Document
Check
Articles

Hybrid Feature Selection and Balancing Data Approach for Improved Software Defect Prediction Febrian, Muhamad Michael; Saputro, Setyo Wahyu; Saragih, Triando Hamonangan; Abadi, Friska; Herteno, Rudy
Indonesian Journal of Electronics, Electromedical Engineering, and Medical Informatics Vol. 7 No. 2 (2025): May
Publisher : Jurusan Teknik Elektromedik, Politeknik Kesehatan Kemenkes Surabaya, Indonesia

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.35882/ijeeemi.v7i2.67

Abstract

Software Defect Prediction (SDP) plays a vital role in identifying defects within software modules. Accurate early detection of software defects can reduce development costs and enhance software reliability. However, SDP remains a significant challenge in the software development lifecycle. This study employs Particle Swarm Optimization (PSO) and addresses several challenges associated with its application, including noisy attributes, high-dimensional data, and imbalanced class distribution. To address these challenges, this study proposed a hybrid filter-based feature selection and class balancing method. The feature selection process incorporates Chi-Square (CS), Correlation-Based Feature Selection (CFS), and Correlation Matrix-Based Feature Selection (CMFS), which have been proven effective in reducing noisy and redundant attributes. Additionally, the Synthetic Minority Over-sampling Technique (SMOTE) is applied to mitigate class imbalance in the dataset. The K-Nearest Neighbors (KNN) algorithm is employed as the classification model due to its simplicity, non-parametric nature, and suitability for handling the feature subsets produced. Performance evaluation is conducted using the Area Under Curve (AUC) metric with a significance threshold of 0.05 to assess classification capability.  The proposed method achieved an AUC of 0.872, demonstrating its effectiveness in enhancing predictive performance. The proposed method was also superior to other combinations such as PSO SMOTE (0.0043), PSO SMOTE CS (0.0091), PSO SMOTE CFS (0.0111), and PSO SMOTE CFS CMFS (0.0007). The findings of this study show that the proposed method significantly enhances the efficiency and accuracy of PSO in software defect prediction tasks. This hybrid strategy demonstrates strong potential as a robust solution for future research and application in predictive software quality assurance.
Analysis of the Effect of Feature Extraction on Sentiment Analysis using BiLSTM: Monkeypox Case Study on X/Twitter Noryasminda; Saragih, Triando Hamonangan; Herteno, Rudy; Faisal, Mohammad Reza; Farmadi, Andi
Indonesian Journal of Electronics, Electromedical Engineering, and Medical Informatics Vol. 7 No. 2 (2025): May
Publisher : Jurusan Teknik Elektromedik, Politeknik Kesehatan Kemenkes Surabaya, Indonesia

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.35882/ijeeemi.v7i2.73

Abstract

The monkeypox outbreak has again become a global concern due to its widespread spread in various countries. Information related to the disease is widely shared through social media, especially Twitter which is a major source of public opinion. However, the complexity of language and the diverse viewpoints of users often pose challenges in accurately analyzing sentiment. Therefore, sentiment analysis of tweets about monkeypox is important to understand public perception and its impact on the dissemination of health information. This research contributes to identifying the most effective word embedding-based feature extraction method for sentiment analysis of health issues on social media. The purpose of this study is to compare the performance of word embedding methods namely Word2Vec, GloVe, and FastText in sentiment analysis of tweets about monkeypox using the BiLSTM model. Data totaling 1511 tweets were collected through a crawling process using the Twitter API. After the data is collected, manual labeling is done into three sentiment categories, namely positive, negative, and neutral. Furthermore, the data is processed through a preprocessing stage which includes data cleaning, case folding, tokenization, stopword removal, and stemming. The evaluation results show that FastText with BiLSTM produces the highest accuracy of 90%, followed by Word2Vec at 89%, and GloVe at 87%. FastText proved to be more effective in reducing classification errors, especially in distinguishing between negative and positive sentiments due to its ability to capture subword information and broader context. These findings suggest that the use of FastText can improve the accuracy of sentiment analysis, especially on health issues that develop on social media, so that it can support data-driven decision making by relevant parties in handling information dissemination. 
Machine Learning Implementation for Sentiment Analysis on X/Twitter: Case Study of Class Of Champions Event in Indonesia Hafizah, Rini; Saragih, Triando Hamonangan; Muliadi, Muliadi; Indriani, Fatma; Mazdadi, Muhammad Itqan
Indonesian Journal of Electronics, Electromedical Engineering, and Medical Informatics Vol. 7 No. 2 (2025): May
Publisher : Jurusan Teknik Elektromedik, Politeknik Kesehatan Kemenkes Surabaya, Indonesia

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.35882/ijeeemi.v7i2.81

Abstract

Sentiment analysis on social media is becoming an important approach in understanding public opinion towards an event. Twitter, as a microblogging platform, generates a large amount of data that can be utilized for this analysis. This study aims to evaluate and compare the performance of three classification algorithms, namely Support Vector Machine (SVM), Random Forest, and Extreme Gradient Boosting (XGBoost), in sentiment analysis related to the Clash of Champions event in Indonesia. To represent the text data, two feature extraction techniques are used, namely Term Frequency-Inverse Document Frequency (TF-IDF) and Bag of Words (BoW). In addition, Synthetic Minority Over-sampling Technique (SMOTE) is applied to handle data imbalance, while model optimization is performed using GridSearchCV. The research dataset consists of 1,000 tweets collected through web scraping, then manually processed and labeled before model training and testing. The results showed that the TF-IDF technique provided superior results compared to BoW. The Random Forest model with TF-IDF achieved the highest accuracy of 91%, while XGBoost with TF-IDF had the highest Area Under the Curve (AUC) of 0.91. The findings confirm that the selection of appropriate feature extraction techniques and algorithms can improve accuracy in sentiment analysis. This study can be applied in public opinion monitoring and data-driven decision-making. Future research can explore word embedding techniques and transformer-based deep learning models to improve semantic understanding and accuracy of sentiment analysis.
Application of Adaboost Algorithm with SMOTE and Optuna Techniques in Sleep Disorder Classification Anshory, Muhammad Naufal; Mazdadi, Muhammad Itqan; Saragih, Triando Hamonangan; Budiman, Irwan; Saputro, Setyo Wahyu
Indonesian Journal of Electronics, Electromedical Engineering, and Medical Informatics Vol. 7 No. 2 (2025): May
Publisher : Jurusan Teknik Elektromedik, Politeknik Kesehatan Kemenkes Surabaya, Indonesia

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.35882/ijeeemi.v7i2.99

Abstract

Data imbalance is a serious challenge in developing machine learning models for sleep disorder classification. When models are trained on an uneven distribution of classes, classification performance for minority classes such as insomnia and sleep apnea is often low. As a result, the overall accuracy may seem elevated, yet the sensitivity to important cases to be weak. Therefore, this research aims to design and develop a robust sleep disorder classification model with the AdaBoost algorithm, with improved performance through the integration of two main approaches, namely data balancing technique utilizing SMOTE and hyperparameter optimization using Optuna. This research contributes by showing that the combination of the two approaches can significantly improve model performance, not only in terms of global accuracy, but also accuracy on previously overlooked minority classes. The dataset utilized is the Sleep Health and Lifestyle Dataset which consists of 374 synthesized data and is divided into three categories: insomnia, sleep apnea, and none. This method stages include data preprocessing, data division using train-test split (80:20), application of SMOTE to balance the class distribution, hyperparameter tuning using Optuna, and model training with the AdaBoost algorithm. Evaluation was performed using classification metrics: accuracy, precision, recall, and F1-score. Results showed that mix of SMOTE and Optuna yielded the best results, accuracy 90.6%, F1-score 0.83871 for insomnia, and 0.81250 for sleep apnea. This performance was consistently superior to scenarios with no SMOTE or no tuning. This confirms the importance of using combination strategies to obtain fair and accurate classification on medical data. Future research is recommended to use real datasets as well as test the capabilities of this research on other models such as XGBoost or LightGBM.
Peningkatan Akurasi Model Boosting pada Prediksi Kesehatan Tidur Menggunakan Optuna Mazdadi, Muhammad Itqan; Saragih, Triando Hamonangan; Budiman, Irwan; Anshory, Muhammad Naufal
Jurnal Informatika Polinema Vol. 12 No. 2 (2026): Vol. 12 No. 2 (2026)
Publisher : UPT P2M State Polytechnic of Malang

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.33795/jip.v12i2.8878

Abstract

Kualitas tidur memiliki peran penting dalam menjaga kesehatan fisik maupun mental, sementara gangguan tidur dapat meningkatkan risiko berbagai penyakit kronis. Perkembangan machine learning membuka peluang untuk melakukan prediksi kesehatan tidur secara lebih akurat melalui pemanfaatan data gaya hidup. Penelitian ini berfokus pada penerapan algoritma boosting, yaitu XGBoost, LightGBM, AdaBoost, dan GradientBoosting, dengan dukungan teknik hyperparameter tuning berbasis Optuna untuk meningkatkan akurasi prediksi. Dataset yang digunakan adalah Sleep Health and Lifestyle Dataset yang memuat variabel demografis, kebiasaan hidup, serta kondisi tidur. Tahapan penelitian meliputi praproses data, pembagian data latih dan uji, pelatihan model, optimasi hyperparameter menggunakan Optuna dengan metode Tree-structured Parzen Estimator (TPE), serta evaluasi model menggunakan metrik akurasi. Hasil eksperimen menunjukkan bahwa tuning dengan Optuna memberikan peningkatan akurasi pada beberapa model, khususnya LightGBM dan AdaBoost, dengan nilai akurasi mencapai 93,3% dan 90,7%. Sementara itu, XGBoost dan GradientBoosting menunjukkan performa stabil dengan akurasi tetap tinggi baik sebelum maupun sesudah tuning. Temuan ini menegaskan bahwa efektivitas tuning bergantung pada karakteristik algoritma yang digunakan. Secara keseluruhan, penelitian ini membuktikan bahwa Optuna dapat menjadi solusi efektif dalam meningkatkan kinerja model boosting untuk prediksi kesehatan tidur. Sebagai arah penelitian lanjutan, disarankan penggunaan metrik evaluasi yang lebih beragam, penerapan teknik penyeimbangan data, serta eksplorasi integrasi dengan metode deep learning untuk memperkaya hasil analisis.
Comparasion Of Weather Classification Methods On Weather Images Using GLCM Features With Random Forest And Catboost Algoritms Noorhafizi, Muhammad; Saragih, Triando Hamonangan; Mazdadi, Muhammad Itqan; Muliadi, Muliadi; Herteno, Rudy; Rozaq, Hasri Awal Akbar
International Journal of Advances in Data and Information Systems Vol. 7 No. 1 (2026): April 2026 - International Journal of Advances in Data and Information Systems
Publisher : Indonesian Scientific Journal

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.59395/ijadis.v7i1.1456

Abstract

Weather image classification is an essential process for improving automated weather information systems. However, most existing studies rely on numerical meteorological data and rarely utilize the textural characteristics embedded in atmospheric imagery. This study addresses that limitation by applying the Gray Level Co-Occurrence Matrix (GLCM) for texture feature extraction combined with Random Forest (RF) and CatBoost algorithms for classification. The dataset, obtained from Kaggle, consists of 1,125 weather images categorized into four classes: cloudy, rain, shine, and sunrise. All images were uniformly normalized and augmented using four rotation angles (0°, 45°, 90°, 135°). GLCM features were extracted with a pixel distance of 1 and gray-level quantization of 8, generating four statistical attributes: contrast, correlation, energy, and homogeneity. Both algorithms were optimized through parameter tuning and evaluated using a 5-fold cross-validation scheme with an 80:20 split ratio. Results show that the Random Forest model (n_estimators = 100, max_depth = 10, random_state = 42) achieved the highest accuracy of 92.43% (±1.12), precision of 92.50%, recall of 92.43%, and F1-score of 92.42%. In comparison, CatBoost (iterations = 100, learning_rate = 0.1, depth = 6) achieved an accuracy of 68.88% (±2.31). The findings demonstrate that GLCM feature extraction combined with Random Forest offers superior stability and accuracy for weather image classification, providing a foundation for efficient and interpretable weather information systems.
Empirical Performance of E2E Frameworks in React-Vue SPAs Using DIA Rezeki, Abdillah; Saputro, Setyo Wahyu; Saragih, Triando Hamonangan; Nugroho, Radityo Adi; Abadi, Friska
International Journal of Advances in Data and Information Systems Vol. 7 No. 1 (2026): April 2026 - International Journal of Advances in Data and Information Systems
Publisher : Indonesian Scientific Journal

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.59395/ijadis.v7i1.1528

Abstract

Modern web applications increasingly adopt Single-Page Application (SPA) architectures to enhance the user experience through client-side rendering and dynamic content loading. However, these characteristics introduce significant challenges for automated end-to-end (E2E) testing, including asynchronous DOM manipulation, complex state management, and timing synchronization issues. This study presents a comprehensive empirical comparison of three prominent E2E testing frameworks—Selenium WebDriver, Cypress, and Playwright—across React and Vue-based SPAs. Using a quantitative experimental approach, 25 standardized test cases were executed 15 times each across Chrome, Firefox, and Edge, for a total of 270 testing sessions. Performance evaluation focused on four key metrics: execution time, success rate, CPU usage, and memory consumption. Results demonstrate that Playwright achieved the fastest execution time (56.25 seconds on React-Chrome), while Selenium exhibited superior resource efficiency with the lowest memory consumption (196.59 MB on Vue-Chrome). The Distance to Ideal Alternative (DIA) multi-criteria decision analysis method identified Playwright-Chrome as optimal for React applications (DIA score: 0.886715) and Selenium-Chrome for Vue applications (DIA score: 0.908237), indicating that framework selection should be context-dependent based on application characteristics and deployment requirements. This research supports the conclusion that no universal "best" testing framework exists, underscoring the importance of evidence-based, application-specific tool selection in software quality assurance.
AdaBoost Classifier untuk Klasifikasi Tanaman Jarak Pagar Triando Hamonangan Saragih; Muliadi Muliadi; Mohammad Reza Faisal; Muhammad Al Ichsan Nur Rizqi Said
Jurnal Komputasi Vol. 9 No. 2 (2021)
Publisher : Jurusan Ilmu Komputer Fakultas MIPA Universitas Lampung

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.23960/komputasi.v9i2.2865

Abstract

Tanaman Jarak Pagar merupakan tanaman multi fungsi yang memiliki banyak kegunaan di kehidupan sehari-hari, baik itu untuk pengobatan, kecantikan hingga pengganti bahan bakar biodiesel. Penyakit yang menyerang tanaman jarak pagar dapat menurunkan kualitas yang dihasilkan jarak pagar. Minimnya pengetahuan petani dan sedikitnya jumlah pakar yang memahami tentang jarak pagar menjadi masalah yang harus diselesaikan. Pengguanaan sistem pakar menjadi solusi yang bisa ditawarkan. AdaBoost Classifier pada sistem pakar dapat digunakan sebagai mengklasifikasikan penyakit tanaman jarak pagar. Hasil yang diperoleh dari penelitian ini yaitu didapat akurasi rata-rata sebesar 50% dan maksimal terbaik sebesar 53,01% pada jumlah fold sebanyak 2. Hasil pada penelitian ini lebih baik dibanding penelitian sebelumnya, tetapi tidak bisa memberikan hasil yang maksimal. Jumlah data tiap kelas menjadi perrmalasahan mengapa hasil pada AdaBoost kurang maksimal dan harus diselesaikan pada penelitian selanjutnya.
Klasifikasi Tanaman Jarak Pagar Menggunakan Algoritme Deep Learning H2O Muhammad Itqan Mazdadi; Rahmat Ramadhani; Triando Hamonangan Saragih; Muhammad Haekal
Jurnal Komputasi Vol. 9 No. 1 (2021)
Publisher : Jurusan Ilmu Komputer Fakultas MIPA Universitas Lampung

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.23960/komputasi.v9i1.2774

Abstract

Tanaman jarak pagar merupakan tanaman multi fungsi yang memiliki banyak manfaat dari daun hingga buah. Tanaman jarak pagar sering digunakan untuk produk kecantikan hingga pengganti biodiesel. Penyakit yang menyerang tanaman jarak pagar dapat mengganggu hasil dari tanaman jarak pagar. Kurangnya pakar dibidang ini dan pengetahuan yang dimiliki petani menyebabkan sesuatu yang buruk. Persoalan ini dapat diselesaikan dengan metode Deep Learning. Metode Deep Learning yang digunakan adalah H2O. H2O digunakan karena dapat memberikan hasil komputasi yang cepat dan bisa memberikan akurasi yang baik. Pada penelitian ini bisa kita lihat bahwa H2O memberikan akurasi rata-rata maksimal sebesar 96,066% dengan parameter uji kombinasi data latih dan data uji 60:40, menggunakan satu layer dan jumlah epoch sebanyak 100. Pada penelitian ini membuktikan bahwa H2O bisa digunakan untuk identifikasi penyakit tanaman jarak pagar.
IMPLEMENTASI CATBOOST DENGAN MENGGUNAKAN HYPER-PARAMETER TUNING BAYESIAN SEARCH UNTUK MEMPREDIKSI PENYAKIT DIABETES Arif Darmawan; Muliadi Muliadi; Dwi Kartini; Triando Hamonangan Saragih; Radityo Adi Nugraha
Jurnal Komputasi Vol. 11 No. 2 (2023)
Publisher : Jurusan Ilmu Komputer Fakultas MIPA Universitas Lampung

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.23960/komputasi.v11i2.13746

Abstract

Diabetes merupakan masalah kesehatan masyarakat dunia dengan prevalensi yang selalu meningkat setiap tahun. Penyakit Diabetes ini perlu didiagnosis sejak dini menggunakan algoritma klasifikasi. Dataset yang digunakan yaitu PIMA Indians Diabetes Database dari Kaggle dengan 768 data dan 8 fitur. Metode pengklasifikasi yang digunakan yaitu Catboost. Klasifikasi Catboost dapat bekerja baik dalam menangani ketidak seimbangan data, namun kinerja algoritma ini masih bisa ditingkatkan lagi. Untuk mengatasi permasalahan tersebut peneliti menggunakan solusi Hyper-parameter tuning. Catboost memiliki beberapa Hyper-parameter yang dapat dikonfigurasi untuk meningkatkan kinerja dari model. Masalah mengidentifikasi nilai yang baik untuk Hyper-parameter disebut Hyper-parameter tuning. Metode Hyper-parameter tuning yang digunakan yaitu Bayesian Search yang kemudian divalidasi menggunakan 10-Fold Cross Validation sebanyak 10 iterasi. Hyper-parameter Catboost yang dikonfigurasi antara lain depth, learning_rate dan Iterations. Pengujian pada Catboost tanpa Hyper-parameter tuning memperoleh nilai presisi sebesar 0,625% dan nilai AUC sebesar 0,868%. Untuk pengujian Catboost dengan Hyper-parameter tuning memperoleh presisi sebesar 0,634 % dan AUC sebesar 0,901%. Menambahkan Hyper-parameter tuning Bayesian Search pada metode klasifikasi Catboost dapat meningkatkan hasil nilai akurasi dan nilai AUC.
Co-Authors AA Sudharmawan, AA Abdul Latief Abadi Abdullayev, Vugar Achmad Rizal Adam Mukharil Bachtiar Adawiyah, Laila Afifa, Ridha Ahmad Rusadi Arrahimi - Universitas Lambung Mangkurat) Ahmad Rusadi Arrahimi - Universitas Lambung Mangkurat) Ahmad Tajali Ahmad Tajali Aida, Nor Ajwa Helisa Al Ghifari, Muhammad Akmal Alamudin, Muhammad Faiq Alfita Rakhmandasari Amelia Aditya Santika Andi Farmadi Andi Farmadi Andi Farmadi Anshari, Muhammad Ridha Anshory, Muhammad Naufal Ansyari, Muhammad Ridho Arif Darmawan Athavale, Vijay Anant Athavale, Vijay Annant Bachtiar, Adam Mukharil Difa Fitria Dina Arifah Diny Melsye Nurul Fajri Diny Melsye Nurul Fajri Dodon Turianto Nugrahadi Dwi Kartini Dwi Kartini, Dwi Dzira Naufia Jawza Erdi, Muhammad Erlianita, Noor Faisal, Mohammad Reza Fatma Indriani Fatma Indriani Febrian, Muhamad Michael Friska Abadi Haekal, Muhammad Haekal, Muhammad Hafizah, Rini Hermiati, Arya Syifa Herteno, Rudy Huynh, Phuoc-Hai Ichwan Dwi Nugraha Indriani, Fatma Irwan Budiman Irwan Budiman Irwan Budiman Irwan Budiman Itqan Mazdadi, Muhammad Ivan Sitohang Jumadi Mabe Parenreng Keswani, Ryan Rhiveldi Lilies Handayani Lumbanraja, Favorisen R M. Khairul Rezki Mafazy, Muhammad Meftah Mariana Dewi Muhamad Fawwaz Akbar Muhammad Al Ichsan Nur Rizqi Said Muhammad Alkaff Muhammad Darmadi Muhammad Fauzan Nafiz Muhammad Haekal Muhammad Haekal Muhammad Ikhwan Rizki Muhammad Itqan Mazdadi Muhammad Mursyidan Amini Muhammad Nadim Mubaarok Muhammad Reza Faisal, Muhammad Reza Muhammad Rofiq Muliadi Muliadi Muliadi Muliadi Muliadi Muliadi Muliadi Muliadi Muliadi Musyaffa, Muhammad Hafizh Noorhafizi, Muhammad Noryasminda Nugraha, Muhammad Amir Nurcahyati, Ica Nurlatifah Amini Okta Muthia Sari Purwoko, Agus Putra, Aditya Maulana Perdana Raditya, Virgi Atha Radityo Adi Nugraha Radityo Adi Nugroho Rahmat Ramadhani Rahmat Ramadhani Rahmatullah, Satrio Wibowo Ramadhani, Rahmat Ratna Septia Devi Regina Reza Faisal, Mohammad Rezeki, Abdillah Rizki, M. Alfi Rozaq, Hasri Akbar Awal Rozaq, Hasri Awal Akbar Rudy Herteno Rudy Herteno Said, Muhammad Al Ichsan Nur Rizqi SALLY LUTFIANI Salsha Farahdiba Setyo Wahyu Saputro Siena, Laifansan Siti Aisyah Solechah Siti Napi'ah Suci Permata Sari Sulastri Norindah Sari Totok Wianto Vivi Nur Wijayaningrum Wahyu Caesarendra Wayan Firdaus Mahmudy Winda Agustina Yanche Kurniawan Mangalik Yasmin Dwi Safitri YILDIZ, Oktay Yusuf Priyo Anggodo Zamzam, Yra Fatria