Claim Missing Document
Check
Articles

Implementasi Reduksi Fitur t-SNE Pada Clustering Gambar Head shape Nematoda Muhammad Rizky Adriansyah; Mohammad Reza Faisal; Abdul Gafur; Radityo Adi Nugroho; Irwan Budiman; Muliadi Muliadi
Jurnal Komputasi Vol. 10 No. 1 (2022)
Publisher : Jurusan Ilmu Komputer Fakultas MIPA Universitas Lampung

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.23960/komputasi.v10i1.2963

Abstract

Pada penelitan ini dilakukan clustering terhadap gambar head shape nematoda, dalam melakukan pengolahan gambar diperlukan metode ekstraksi fitur untuk menemukan informasi penting dari gambar yang akan diolah, salah satu esktraksi fitur yang bisa digunakan adalah wavelet. Setelah gambar melewati ekstraksi fitur dihasilkan sebanyak 5624 fitur, dengan fitur sebanyak ini dapat mengakibatkan waktu komputasi yang lama. Oleh sebab itu perlu dilakukan reduksi fitur untuk mengurangi jumlah fitur yang awalnya 5624 fitur menjadi 2 atau 3 fitur saja, salah satu metode reduksi fitur terbaru yang bisa digunakan adalah t-SNE. Pada penelitian ini dilakukan perbandingan hasil kualitas cluster antara yang menggunakan reduksi fitur dengan yang tidak. Hasil Silhouette Index   yang didapatkan tanpa reduksi fitur adalah 0.046 dan setelah menggunakan reduksi fitur t-SNE terjadi peningkatan yang cukup signifikan menjadi 0.418.
Android Malware Detection with Hybrid Feature Selection and Bayesian Optimization Fadhillah, Muhammad Alif; Saputro, Setyo Wahyu; Muliadi, Muliadi; Faisal, Mohammad Reza; Nugroho, Radityo Adi
International Journal of Advances in Data and Information Systems Vol. 7 No. 1 (2026): April 2026 - International Journal of Advances in Data and Information Systems
Publisher : Indonesian Scientific Journal

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.59395/ijadis.v7i1.1526

Abstract

The increasing dimensionality of Android application features poses significant challenges for accurate and efficient malware detection. This study proposes a hybrid feature selection framework that combines Minimum Redundancy Maximum Relevance (mRMR) and correlation filtering to optimize classification performance on the Drebin-215 dataset. A selected configuration of 175 features with a correlation threshold of 0.7 was evaluated using five classifiers: LSTM, Support Vector Machine (SVM), Random Forest, K-Nearest Neighbors (KNN), and XGBoost. The experimental results show that dimensionality reduction improves classification stability and overall predictive performance. SVM exhibits the most notable improvement, with accuracy increasing from 63.05% without feature selection to 98.57% after applying the proposed framework. LSTM achieves 98.57% accuracy with an AUC of 99.86%, while Random Forest, KNN, and XGBoost consistently achieve accuracy above 97%. In addition to performance enhancement, the hybrid feature selection approach substantially improves computational efficiency. SVM training time decreases from 770.75 seconds to 155.88 seconds, and testing time is reduced from 15.581 seconds to 0.3824 seconds. KNN testing time also decreases from 1.623 seconds to 0.4595 seconds..
Enhancing Software Defect Prediction through Hybrid Multi-Filter Feature Selection and Imbalance Handling Muhammad Khalid Maulana; Setyo Wahyu Saputro; Mohammad Reza Faisal; Radityo Adi Nugroho; As’ary Ramadhan
Journal of Computing Theories and Applications Vol. 3 No. 4 (2026): JCTA 3(4) 2026
Publisher : Universitas Dian Nuswantoro

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.62411/jcta.15943

Abstract

Software Defect Prediction (SDP) aims to identify defective modules early in the software development lifecycle to improve software quality and reduce maintenance costs. However, SDP datasets commonly suffer from high dimensionality, feature redundancy, and class imbalance, which can degrade model performance and stability. This study proposes a hybrid feature selection framework to address these challenges and enhance prediction performance. The proposed approach integrates Combined Correlation and Mutual Information (CONMI), which combines the Pearson Correlation Coefficient (PCC) and Mutual Information (MI) to capture both linear and nonlinear feature relevance. The selected features are further refined through Top-K selection, correlation-based filtering to reduce multicollinearity, and Backward Elimination (BE) to obtain an optimal feature subset. To address class imbalance, SMOTE-Tomek is applied by combining over-sampling and data cleaning techniques. Experiments are conducted on twelve NASA MDP datasets using Logistic Regression (LR) and Naïve Bayes (NB) classifiers. The results show that the proposed framework consistently achieves the best performance, with Logistic Regression combined with SMOTE-Tomek obtaining the highest average AUC of 0.7923 ± 0.0714, while NB achieves 0.7554 ± 0.0580. Statistical analysis using a paired t-test indicates that the proposed method significantly outperforms MI+SMOTE-Tomek and BE+SMOTE-Tomek for Logistic Regression, whereas no significant differences are observed for NB. In addition to improving overall classification performance (AUC), the proposed approach also enhances minority class detection, as reflected in improved Recall and F1-score. Overall, the proposed hybrid framework provides an effective and reliable solution for software defect prediction, particularly for high-dimensional and imbalanced datasets.
Quantifying the Impact of Text Preprocessing on IndoBERT Fine-Tuning for Indonesian Informal Culinary Sentiment Analysis Rahmat Budianoor; Setyo Wahyu Saputro; Friska Abadi; Radityo Adi Nugroho; Andi Farmadi
Journal of Computing Theories and Applications Vol. 3 No. 4 (2026): JCTA 3(4) 2026
Publisher : Universitas Dian Nuswantoro

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.62411/jcta.15980

Abstract

Indonesian culinary comments on social media platforms such as Instagram are characterized by informal spelling, regional language mixing, slang expressions, and emojis, posing substantial challenges for automated sentiment classification. While IndoBERT has demonstrated strong performance across Indonesian natural language processing tasks, the contribution of individual preprocessing components to fine-tuning performance on informal text remains underexplored, particularly in the culinary domain. This study addresses this gap by conducting a systematic preprocessing ablation study on IndoBERT-Base fine-tuning for Indonesian culinary sentiment classification, accompanied by a comparative evaluation against Naive Bayes with TF-IDF, SVM with TF-IDF, and BiLSTM as representative baselines. A dataset of 3,500 manually labeled Instagram culinary comments across three sentiment classes was used, with a stratified 80/10/10 split. Six preprocessing variants were evaluated under identical experimental conditions to isolate the contribution of each component. The results show that slang normalization is the most impactful single preprocessing step, yielding a macro F1-score gain of +0.0609 over the no-preprocessing baseline, while the full pipeline achieves an accuracy of 0.8800 and a macro F1-score of 0.8465. IndoBERT-Base with the full pipeline outperforms all baselines across all evaluation metrics. Per-class analysis reveals that the negative class achieves the lowest F1-score of 0.7600, with sarcastic expressions and Banjar regional vocabulary identified as primary sources of misclassification. These findings indicate that preprocessing decisions have a measurable and non-uniform effect on IndoBERT fine-tuning performance. In this study, slang normalization provides the most substantial individual contribution in bridging the vocabulary gap between informal user-generated text and the model’s pre-training distribution.
Optimized multi correlation-based feature selection in software defect prediction Muhammad Nabil Muyassar Rahman; Radityo Adi Nugroho; Mohammad Reza Faisal; Friska Abadi; Rudy Herteno
TELKOMNIKA (Telecommunication Computing Electronics and Control) Vol 22, No 3: June 2024
Publisher : Universitas Ahmad Dahlan

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.12928/telkomnika.v22i3.25793

Abstract

In software defect prediction, noisy attributes and high-dimensional data remain to be a critical challenge. This paper introduces a novel approach known as multi correlation-based feature selection (MCFS), which seeks to address these challenges. MCFS integrates two feature selection techniques, namely correlation-based feature selection (CFS) and correlation matrixbased feature selection (CMFS), intending to reduce data dimensionality and eliminate noisy attributes. To accomplish this, CFS and CMFS are applied independently to filter the datasets, and a weighted average of their outcomes is computed to determine the optimal feature selection. This approach not only reduces data dimensionality but also mitigates the impact of noisy attributes. To further enhance predictive performance, this paper leverages the particle swarm optimization (PSO) algorithm as a feature selection mechanism, specifically targeting improvements in the area under the curve (AUC). The evaluation of the proposed method is conducted on 12 benchmark datasets sourced from the NASA metrics data program (MDP) corpus, renowned for their noisy attributes, high dimensionality, and imbalanced class records. The research findings demonstrate that MCFS outperforms CFS and CMFS, yielding an average AUC value of 0.891, thereby emphasizing it is efficacy in advancing classification performance in the context of software defect prediction using k-nearest neighbors (KNN) classification.
The impact of software metrics in NASA metric data program dataset modules for software defect prediction Adinda Ayu Puspita Ramadhani; Radityo Adi Nugroho; Mohammad Reza Faisal; Friska Abadi; Rudy Herteno
TELKOMNIKA (Telecommunication Computing Electronics and Control) Vol 22, No 4: August 2024
Publisher : Universitas Ahmad Dahlan

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.12928/telkomnika.v22i4.25787

Abstract

This paper discusses software metrics and their impact on software defect prediction values in the NASA metric data program (MDP) dataset. The NASA MDP dataset consists of four categories of software metrics: halstead, McCabe, LoC, and misc. However, there is no study showing which metrics participate in increasing the area under the curve (AUC) value of the NASA MDP dataset. This study utilizes 12 modules from the NASA MDP dataset, where these 12 modules are being tested into 14 relationships of software metrics derived from the four existing metric categories. Subsequently, classification is performed using the k-nearest neighbor (kNN) method. The research concludes that software metrics have a significant impact on the AUC value, with the LoC+McCabe+misc metrics relationship influencing the improvement of the AUC value. However, the metrics relationship that has the most impact on achieving less optimal AUC values is McCabe. Halstead metric also plays a role in decreasing the performance of other metrics.
Co-Authors Abdul Gafur Adi Mu'Ammar, Rifqi Adin Nofiyanto, Adin Adinda Ayu Puspita Ramadhani Ahmad Bahroini Ahmad Juhdi Ahmad Rusadi Aida, Nor Akhtar, Zarif Bin Alamudin, Muhammad Faiq Andi Farmadi Andi Farmadi Andi Farmadi Angga Maulana Akbar Arie Sapta Nugraha Arie Sapta Nugraha Aryanti, Agustia Kuspita Athavale, Vijay Anant Aylwin Al Rasyid Bayu Hadi Sudrajat Dendy Fadhel Adhipratama Dendy Deni Kurnia Dike Bayu Magfira, Dike Bayu Dodon Turianto Nugrahadi Dwi Kartini Dwi Kartini, Dwi Efendi Mohtar Emma Andini Erdi, Muhammad Fadhillah, Muhammad Alif Faisal, Mohammad Reza Fatma Indriani Fauzan Luthfi, Achmad Fenny Winda Rahayu Fhadilla Muhammad Friska Abadi Friska Abadi Hanif Rahardian Herteno, Rudy Irwan Budiman Irwan Budiman Irwan Budiman Itqan Mazdadi, Muhammad Ivan Sitohang Maya Yusida Muhammad Angga Wiratama Muhammad Azmi Adhani Muhammad Fikri Muhammad Ikhwanul Hakim Muhammad Itqan Mazdadi Muhammad Khalid Maulana Muhammad Latief Saputra Muhammad Nabil Muyassar Rahman Muhammad Noor Muhammad Reza Faisal, Muhammad Reza Muhammad Rizky Adriansyah Muhammad Rusli Muhammad Syahriani Noor Basya Basya Muhammad Zaien Muliadi Muliadi Muliadi Aziz Muliadi Muliadi Muliadi Muliadi Nur Hidayatullah, Wildan Nur Ridha Apriyanti Oni Soesanto Pratama, Muhammad Yoga Adha Putri, Nitami Lestari Rahmat Budianoor Rahmat Ramadhani Raidra Zeniananto Ramadhan, As'ary Reina Alya Rahma Reza Faisal, Mohammad Rezeki, Abdillah Riadi, Putri Agustina Rinaldi Rizal, Muhammad Nur Rizky Ananda, Muhammad Rozaq, Hasri Akbar Awal Rudy Herteno Rudy Herteno Rudy Herteno Salsha Farahdiba Saragih, Triando Hamonangan Sarah Monika Nooralifa Septiadi Marwan Annahar Setyo Wahyu Saputro Siena, Laifansan Suci Permata Sari Suryadi, Mulia Kevin Sutan Takdir Alam Wahyu Caesarendra Wahyu Ramadansyah Zaini Abdan