Muhamad Rosdiana
Universitas Pamulang, Tangerang Selatan

Published : 2 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 2 Documents
Search

Analisis Pengaruh Chi-Square Feature Selection terhadap Kinerja Random Forest dan XGBoost dalam Prediksi Konversi Pengunjung Website Muhamad Rosdiana; Teti Desyani; Perani Rosyani
Building of Informatics, Technology and Science (BITS) Vol 8 No 1 (2026): June 2026
Publisher : Forum Kerjasama Pendidikan Tinggi

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47065/bits.v8i1.9662

Abstract

The increasing number of visitors to e-commerce websites is not always accompanied by a corresponding increase in purchase transactions, making it difficult for companies to identify visitors with high conversion potential. In addition, using all available attributes may increase model complexity without necessarily improving predictive performance. This study analyzes the impact of Chi-Square Feature Selection on the performance of Random Forest and Extreme Gradient Boosting (XGBoost) in predicting website visitor conversion. The study uses the Online Shoppers Purchasing Intention dataset consisting of 12,330 instances with 17 predictor attributes and one target attribute. The research process includes exploratory data analysis, preprocessing, Chi-Square-based feature selection, classification model development using Random Forest and XGBoost, and evaluation using Accuracy, Precision, Recall, F1-Score, Matthews Correlation Coefficient (MCC), and Area Under the Curve (AUC-ROC). Four experimental scenarios were evaluated: all features (baseline), Top-15, Top-10, and Top-5 selected features. The results show that the baseline model using all features achieved the best overall performance. The Random Forest baseline model obtained an Accuracy of 90.05%, Precision of 73.94%, F1-Score of 63.13%, and MCC of 0.5835, while the XGBoost baseline model achieved the highest AUC-ROC of 0.9271. Furthermore, PageValues, BounceRates, ExitRates, ProductRelated_Duration, and ProductRelated were identified as the most influential features affecting visitor conversion. The main contribution of this study is providing empirical evidence that Chi-Square Feature Selection is more effective in reducing feature complexity and identifying relevant attributes than improving classification performance on the Online Shoppers Purchasing Intention dataset, offering practical guidance for feature selection strategies in machine learning-based website conversion prediction
Klasifikasi Batu Permata Berbasis Citra Menggunakan Convolutional Neural Network Perani Rosyani; Oke Hariansyah; Yuda Permadi; Muhamad Rosdiana; Nanang Nanang
Journal of Information System Research (JOSH) Vol 7 No 2 (2026): January 2026
Publisher : Forum Kerjasama Pendidikan Tinggi (FKPT)

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47065/josh.v7i2.9101

Abstract

Manual gemstone identification still faces several limitations, such as subjective assessment and strong dependence on expert experience, which may lead to misclassification, particularly for gemstones with similar visual characteristics. This study aims to apply a Convolutional Neural Network (CNN) for automatic visual-based gemstone image classification using a limited dataset. The dataset consists of three gemstone classes, namely Alexandrite, Almandine, and Amazonite, with a balanced class distribution. Image preprocessing includes image resizing, pixel value normalization, and data augmentation to increase data variability. The proposed CNN model is a custom architecture composed of three convolutional layers with ReLU activation, followed by max pooling, a fully connected layer with dropout, and a Softmax output layer. Model performance is evaluated using a confusion matrix and classification metrics, including accuracy, precision, recall, and F1-score. Experimental results show that the CNN model achieves a testing accuracy of 93.33% on the limited test dataset with relatively balanced performance across classes. However, analysis of the training and validation curves indicates the presence of overfitting, suggesting that the model’s generalization capability to unseen data remains limited. These findings highlight that the achieved accuracy is conditional on the specific and constrained dataset used. Therefore, future work is recommended to expand dataset size and diversity, apply more comprehensive data augmentation strategies, and explore transfer learning approaches to improve model stability and generalization performance.