Claim Missing Document
Check
Articles

Comparison of LASSO, Ridge, and Elastic Net Regularization with Balanced Bagging Classifier Nisrina Az-Zahra, Putri; Sadik, Kusman; Suhaeni, Cici; Mohamad Soleh, Agus
Parameter: Jurnal Matematika, Statistika dan Terapannya Vol 4 No 2 (2025): Parameter: Jurnal Matematika, Statistika dan Terapannya
Publisher : Jurusan Matematika FMIPA Universitas Pattimura

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.30598/parameterv4i2pp287-296

Abstract

Predicting Drug-Induced Autoimmunity (DIA) is crucial in pharmaceutical safety assessment, as early identification of compounds with autoimmune risk can prevent adverse drug reactions and improve patient outcomes. Classification analysis often faces challenges when the number of predictor variables exceeds the number of observations or when high correlations among predictors lead to multicollinearity and overfitting. Regularization methods, such as Ridge Regression, Least Absolute Shrinkage and Selection Operator (LASSO), and Elastic-Net, help stabilize parameter estimation and improve model interpretability. This study focuses on building a binary classification model to predict the risk of DIA using 196 molecular descriptors derived from chemical compound structures. To address class imbalance in the response variable, the Balanced Bagging Classifier (BBC) is combined with regularized logistic regression models. Elastic Net + BBC outperforms other models with the highest accuracy (0.825), followed closely by LASSO + BBC and Ridge + BBC (both 0.816). This integration not only improves classification accuracy but also enhances generalization and the reliable detection of minority class instances, supporting the early identification of autoimmune risks in drug discovery.
EVALUATING RANDOM FOREST AND XGBOOST FOR BANK CUSTOMER CHURN PREDICTION ON IMBALANCED DATA USING SMOTE AND SMOTE-ENN Andespa, Reyuli; Sadik, Kusman; Suhaeni, Cici; Soleh, Agus M
MEDIA STATISTIKA Vol 18, No 1 (2025): Media Statistika
Publisher : Department of Statistics, Faculty of Science and Mathematics, Universitas Diponegoro

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.14710/medstat.18.1.25-36

Abstract

The banking industry faces significant challenges in retaining customers, as churn can critically affect both revenue and reputation. This study introduces a robust churn prediction framework by comparing the performance of XGBoost and Random Forest algorithms under imbalanced data conditions. The novelty of this research lies in integrating the SMOTE and SMOTE-ENN techniques with machine learning algorithms to enhance model performance and reliability on highly imbalanced datasets. Unlike conventional approaches that rely solely on oversampling or undersampling, this study demonstrates that the hybrid combination of XGBoost and SMOTE provides superior predictive accuracy, stability, and efficiency. Hyperparameter optimization using GridSearchCV was conducted to identify the most effective parameter configurations for both algorithms. Model performance was evaluated using the F1-Score and Area Under the Curve (AUC). The results indicate that XGBoost with SMOTE achieved the best performance, with an F1-Score of 0.8730 and an AUC of 0.9828, showing an optimal balance between precision and recall. Feature importance analysis identified Months_Inactive_12_mon, Total_Trans_Amt, and Total_Relationship_Count as the most influential predictors. Overall, this approach outperforms traditional resampling and modeling techniques, providing practical insights for data-driven customer retention strategies in the banking industry.
Deep Learning Image Classification Rontgen Dada pada Kasus Covid-19 Menggunakan Algoritma Convolutional Neural Network Susanti, Leni Anggraini; Soleh, Agus Mohamad; Sartono, Bagus
Jurnal Teknologi Informasi dan Ilmu Komputer Vol 10 No 5: Oktober 2023
Publisher : Fakultas Ilmu Komputer, Universitas Brawijaya

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.25126/jtiik.2023107142

Abstract

Penelitian ini mengusulkan penggunaan Convolutional Neural Network (CNN) dengan arsitektur VGGNet-19 dan ResNet-50 untuk diagnosis COVID-19 melalui analisis citra rontgen dada. Modifikasi dilakukan dengan membandingkan nilai regularisasi dropout 50% dan 80% untuk kedua arsitektur dan mengubah jumlah lapisan klasfikasi menjadi 4 kelas. Selanjutnya, kinerja model dibandingkan berdasarkan ukuran dataset. Dataset terdiri dari 21165 citra, dengan pembagian 10% sebagai data uji dan 90% data dibagi menjadi data latih (80%) dan data validasi (20%). Kinerja model dievaluasi menggunakan metode validasi silang berulang 5 kali lipat. Proses pelatihan menggunakan learning rate 0.0001, optimasi stochastic gradient descent (SGD), dan sepuluh iterasi. Hasil penelitian menunjukkan bahwa penambahan lapisan dropout dengan peluang 50% untuk kedua arsitektur secara efektif mengatasi overfitting dan meningkatkan performa model. Ditemukan bahwa kinerja yang lebih baik dicapai pada ukuran kumpulan data lebih besar dan memberikan peningkatan signifikan pada kinerja model. Hasil klasifikasi menunjukkan arsitektur ResNet-50 mencapai akurasi rata-rata 94.4%, recall rata-rata 94.1%, presisi rata-rata 95.5%, spesifisitas rata-rata 97% dan F1-score rata-rata 94.8%. Sedangkan arsitektur VGGNet-19 mencapai akurasi rata-rata 91%, recall rata-rata 89%, presisi rata-rata 95.0%, spesifisitas rata-rata 96.8% dan F1-score rata-rata 92.7%. Pemanfaatan model ini dapat membantu mengidentifikasi penyebab kematian pasien dan memberikan informasi yang berharga bagi pengambilan keputusan medis dan epidemiologi.   Abstract This research proposes using a Convolutional Neural Network (CNN) with VGGNet-19 and ResNet-50 architectures for COVID-19 diagnosis through chest X-ray image analysis. Modifications were made by comparing the dropout regularization values of 50% and 80% for both architectures and altering the number of classification layers to 4 classes. Furthermore, the model's performance was compared based on dataset size. The dataset comprised 21,165 images, with a division of 10% for testing and 90% divided into training data (80%) and validation data (20%). The model's performance was evaluated using the 5-fold repeat cross-validation method. The training process employed a learning rate of 0.0001, stochastic gradient descent (SGD) optimization, and ten iterations. The study's results indicate that adding dropout layers with a 50% probability for both architectures effectively addressed overfitting and improved the model's performance. It was found that better performance was achieved with larger dataset sizes. The classification results indicate the ResNet-50 architecture achieved an average accuracy of 94.4%, average recall of 94.1%, average precision of 95.5%, average specificity of 97%, and average F1-score of 94.8%. Meanwhile, the VGGNet-19 architecture achieved an average accuracy of 91%, an average recall of 89%, average precision of 95.0%, average specificity of 96.8%, and an average F1-score of 92.7%. Utilizing these models can assist in identifying the causes of patient mortality and offer valuable information for medical and epidemiological decision-making.
Classification Modeling with RNN-based, Random Forest, and XGBoost for Imbalanced Data: A Case of Early Crash Detection in ASEAN-5 Stock Markets Siswara, Deri; M. Soleh, Agus; Hamim Wigena, Aji
Scientific Journal of Informatics Vol. 11 No. 3: August 2024
Publisher : Universitas Negeri Semarang

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.15294/sji.v11i3.4067

Abstract

Purpose: This research aims to evaluate the performance of several Recurrent Neural Network (RNN) architectures, including Simple RNN, Gated Recurrent Units (GRU), and Long Short-Term Memory (LSTM), compared to classic algorithms such as Random Forest and XGBoost, in building classification models for early crash detection in the ASEAN-5 stock markets. Methods: The study examines imbalanced data, which is expected due to the rarity of market crashes. It analyzes daily data from 2010 to 2023 across the major stock markets of the ASEAN-5 countries: Indonesia, Malaysia, Singapore, Thailand, and the Philippines. A market crash is the target variable when the primary stock price indices fall below the Value at Risk (VaR) thresholds of 5%, 2.5%, and 1%. Predictors include technical indicators from major local and global markets and commodity markets. The study incorporates 213 predictors with their respective lags (5, 10, 15, 22, 50, 200) and uses a time step of 7, expanding the total number of predictors to 1,491. The challenge of data imbalance is addressed with SMOTE-ENN. Model performance is evaluated using the false alarm rate, hit rate, balanced accuracy, and the precision-recall curve (PRC) score. Result: The results indicate that all RNN-based architectures outperform Random Forest and XGBoost. Among the various RNN architectures, Simple RNN is the most superior, primarily due to its simple data characteristics and focus on short-term information. Novelty: This study enhances and extends the range of phenomena observed in previous studies by incorporating variables such as different geographical zones and periods and methodological adjustments.
Performance of Ensemble Learning in Diabetic Retinopathy Disease Classification Nurizki, Anisa; Fitrianto, Anwar; Mohamad Soleh, Agus
Scientific Journal of Informatics Vol. 11 No. 2: May 2024
Publisher : Universitas Negeri Semarang

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.15294/sji.v11i2.4725

Abstract

Purpose: This study explores diabetic retinopathy (DR), a complication of diabetes leading to blindness, emphasizing early diagnostic interventions. Leveraging Macular OCT scan data, it aims to optimize prevention strategies through tree-based ensemble learning. Methods: Data from RSKM Eye Center Padang (October-December 2022) were categorized into four scenarios based on physician certificates: Negative & non-diagnostic DR versus Positive DR, Negative versus Positive DR, Non-Diagnosis versus Positive DR, and Negative DR versus non-Diagnosis versus Positive DR. The suitability of each scenario for ensemble learning was assessed. Class imbalance was addressed with SMOTE, while potential underfitting in random forest models was investigated. Models (RF, ET, XGBoost, DRF) were compared based on accuracy, precision, recall, and speed. Results: Tree-based ensemble learning effectively classifies DR, with RF performing exceptionally well (80% recall, 78.15% precision). ET demonstrates superior speed. Scenario III, encompassing positive and undiagnosed DR, emerges as optimal, with the highest recall and precision values. These findings underscore the practical utility of tree-based ensemble learning in DR classification, notably in Scenario III. Novelty: This research distinguishes itself with its unique approach to validating tree-based ensemble learning for DR classification. This validation was accomplished using Macular OCT data and physician certificates, with ETDRS scores demonstrating promising classification capabilities.
Evaluating Ensemble Learning Techniques for Class Imbalance in Machine Learning: A Comparative Analysis of Balanced Random Forest, SMOTE-RF, SMOTEBoost, and RUSBoost Fulazzaky, Tahira; Saefuddin, Asep; Soleh, Agus Mohamad
Scientific Journal of Informatics Vol. 11 No. 4: November 2024
Publisher : Universitas Negeri Semarang

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.15294/sji.v11i4.15937

Abstract

Purpose: This research aims to identify the optimal ensemble learning method for mitigating class imbalance in datasets utilizing various advanced techniques which include balanced random forest (BRF), SMOTE-random forest (SMOTE-RF), RUSBoost, and SMOTEBoost. The methods were systematically evaluated against conventional algorithms, including random forest and AdaBoost, across heterogeneous datasets with varying class imbalance ratios. Methods: This study utilized 13 secondary datasets from diverse sources, each with binary class outputs. The datasets exhibited varying degrees of class imbalance, offering scenarios to assess the effectiveness of ensemble learning techniques and traditional machine learning approaches in managing class imbalance issues. Study data were split into training (80%) and testing (20%), with stratified sampling applied to maintain consistent class proportions across both sets. Each method underwent hyperparameter optimization with distinct settings with repetition over 10 iterations. The optimal method was evaluated based on balanced accuracy, recall, and computation time. Result: Based on the evaluation, the BRF method exhibited the highest performance in balanced accuracy and recall when compared to SMOTE-RF, RUSBoost, SMOTEBoost, random forest, and AdaBoost. Conversely, the classical random forest method outperformed other techniques in terms of computational efficiency. Novelty: This study presents an innovative analysis of advanced ensemble learning techniques, including BRF, SMOTE-random forest, SMOTEBoost, and RUSBoost, which demonstrate significant effectiveness in addressing class imbalance across various datasets. By systematically optimizing hyperparameters and applying stratified sampling, this research produces findings that redefine the benchmarks of balanced accuracy, recall and computational efficiency in machine learning.
A Hybrid Sampling Approach for Handling Data Imbalance in Ensemble Learning Algorithms Astari, Reka Agustia; Sumertajaya, I Made; Soleh, Agus Mohamad
Scientific Journal of Informatics Vol. 12 No. 2: May 2025
Publisher : Universitas Negeri Semarang

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.15294/sji.v12i2.19163

Abstract

Purpose: This research aims to address the methodological challenges posed by imbalanced data in classification tasks, where minority classes are severely underrepresented, often leading to biased model performance. It evaluates the effectiveness of hybrid sampling techniques specifically, the Synthetic Minority Oversampling Technique combined with Neighborhood Cleaning Rule (SMOTE-NCL) and with Edited Nearest Neighbors (SMOTE-ENN) in improving the predictive performance of ensemble classifiers, namely Double Random Forest (DRF) and Extremely Randomized Trees (ET), with a focus on enhancing minority class detection. Methods: A total of eighteen simulated scenarios were developed by varying class imbalance ratios, sample sizes, and feature correlation levels. In addition, empirical data from the 2023 National Socioeconomic Survey (SUSENAS) in Riau Province were employed. The data were partitioned using stratified random sampling (80% training, 20% testing). Models were trained with and without hybrid sampling and optimized through grid search. Their performance was evaluated over 100 iterations using balanced accuracy, sensitivity, and G-mean. Feature importance was interpreted using Shapley Additive Explanations (SHAP). Results: DRF combined with SMOTE-NCL consistently outperformed all other models, achieving 87.56% balanced accuracy, 82.17% sensitivity, and 86.75% G-mean in the most extreme simulation scenario. On the empirical dataset, the model achieved 76.37% balanced accuracy and 75.49% G-mean. Novelty: This study introduces a novel integration of hybrid sampling techniques and ensemble learning within an interpretable machine learning framework, providing a robust solution for poverty classification in imbalanced datasets.
Comparison of Ensemble Forest-Based Methods Performance for Imbalanced Data Classification Hasnataeni, Yunia; Saefuddin, Asep; Soleh, Agus Mohamad
Scientific Journal of Informatics Vol. 12 No. 2: May 2025
Publisher : Universitas Negeri Semarang

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.15294/sji.v12i2.24269

Abstract

Purpose: Classification of imbalanced data presents a major challenge in meteorological studies, particularly in rainfall classification where extreme events occur infrequently. This research addresses the issue by evaluating ensemble learning models in handling imbalanced rainfall data in Bogor Regency, aiming to improve classification performance and model reliability for hydrometeorological risk mitigation. Methods: Four ensemble methods: RF, RoF, DRF, and RoDRF were applied to rainfall classification using three resampling techniques: SMOTE, RUS, and SMOTE-RUS-NC. The data underwent preprocessing, stratified splitting, resampling, and 5-fold cross-validation. Performance was evaluated over 100 iterations using accuracy, precision, recall, and F1-score. Result: The combination of DRF with SMOTE-RUS-NC yielded the most balanced results between accuracy (0.989) and computation time (107.28 seconds), while RoDRF with SMOTE achieved the highest overall performance with an accuracy of 0.991 but required a longer computation time (149.30 seconds). Feature importance analysis identified average humidity, maximum temperature, and minimum temperature as the most influential predictors of extreme rainfall. Novelty: This research contributes a comprehensive comparison of ensemble forest-based methods for imbalanced rainfall data, revealing DRF-SMOTE as an optimal trade-off between performance and efficiency. The findings contribute to improved rainfall classification models and offer practical insight for disaster mitigation planning and resource management in tropical regions.
Land Use Change Modelling Using Logistic Regression, Random Forest and Additive Logistic Regression in Kubu Raya Regency, West Kalimantan Pradana, Alfa Nugraha; Djuraidah, Anik; Soleh, Agus Mohamad
Forum Geografi Vol 37, No 2 (2023): December 2023
Publisher : Universitas Muhammadiyah Surakarta

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.23917/forgeo.v37i2.23270

Abstract

Kubu Raya Regency is a regency in the province of West Kalimantan which has a wetland ecosystem including a high-density swamp or peatland ecosystem along with an extensive area of mangroves. The function of wetland ecosystems is essential for fauna, as a source of livelihood for the surrounding community and as storage reservoir for carbon stocks. Most of the land in Kubu Raya Regency is peatland. As a consequence, peat has long been used for agriculture and as a source of livelihood for the community. Along with the vast area of peat, the regency also has a potential high risk of peat fires. This study aims to predict land use changes in Kubu Raya Regency using three statistical machine learning models, specifically Logistic Regression (LR), Random Forest (RF) and Additive Logistic Regression (ALR). Land cover map data were acquired from the Ministry of Environment and Forestry and subsequently reclassified into six types of land cover at a resolution of 100 m. The land cover data were employed to classify land use or land cover class for the Kubu Raya regency, for the years 2009, 2015 and 2020. Based on model performance, RF provides greater accuracy and F1 score as opposed to LR and ALR. The outcome of this study is expected to provide knowledge and recommendations that may aid in developing future sustainable development planning and management for Kubu Raya Regency.
Support vector machine performance: simulation and rice phenology application Muradi, Hengki; Saefuddin, Asep; Sumertajaya, I Made; Soleh, Agus Mohamad; Domiri, Dede Dirgahayu
IAES International Journal of Artificial Intelligence (IJ-AI) Vol 14, No 6: December 2025
Publisher : Institute of Advanced Engineering and Science

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.11591/ijai.v14.i6.pp4878-4890

Abstract

In the case of classification, model accuracy is expected to result in correct predictions. This study aims to analyze the performance of two kinds of support vector machine (SVM) methods: the support vector machine one versus one (SVM OvO) method and the generalized multiclass support vector machine (GenSVM) method. This method will compare to the generalized linear model, namely the multinomial logistic regression (MLR) method. Simulations were conducted using SVM OvO and GenSVM methods to get an overview of the parameters affecting both methods' performance. Furthermore, the three classification methods are implemented in the case of modelling the rice phenology and tested for performance. Simulation results show that, however, the SVM OvO and GenSVM machine learning methods are sensitive to the choice of model parameters. The empirical study results show that the SVM OvO and GenSVM methods can produce satisfactory model accuracy and are comparable to the MLR method. The best rice phenology model accuracy was obtained from the SVM OvO model, where 79.20 ± 0.21 overall accuracy and 70.69 ± 0.29 kappa were obtained. This research can be continued by handling samples, especially when class members are a minority, and can also add random effects to the SVM model.