Claim Missing Document
Check
Articles

Found 3 Documents
Search

Hybrid LBFA-Based Feature Selection for Improving Machine Learning Classification Performance in Heart Disease Prediction Hana Azizah; Eni Sumarminingsih; Adji Achmad Rinaldo Fernandes
UNP Journal of Statistics and Data Science Vol. 4 No. 2 (2026): UNP Journal of Statistics and Data Science
Publisher : Departemen Statistika Universitas Negeri Padang

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.24036/ujsds/vol4-iss2/478

Abstract

Feature selection and feature engineering are essential steps in developing accurate machine learning models, particularly when dealing with imbalanced datasets and redundant variables. However, many feature augmentation methods are often applied without a consistent preprocessing strategy, which can reduce model reliability and increase the risk of information leakage. To overcome this issue, this study proposes a hybrid classification framework that combines CatBoost-based feature selection with two feature augmentation techniques: LOGIT transformation and Log Density Ratio (LDR). A structured preprocessing pipeline was designed to ensure consistency throughout the modeling process. One-hot encoding was applied for the LOGIT transformation, while numerical standardization was used for LDR estimation. The generated features were then integrated with the selected original variables to produce richer feature representations for classification. The proposed framework was evaluated using the Heart Disease dataset with three gradient boosting algorithms, namely LightGBM, XGBoost, and CatBoost. Model performance was assessed using accuracy, precision, sensitivity, specificity, and F1-score. The results show that the proposed approach consistently improved classification performance across all models. Among the tested models, LightGBM combined with LOGIT and LDR achieved the best performance, obtaining an accuracy of 0.9618, precision of 0.9485, sensitivity of 0.9620, specificity of 0.9625, and F1-score of 0.9552. These findings suggest that combining feature selection with structured feature augmentation can significantly improve predictive performance in imbalanced classification tasks
Modelling Geographically Weighted Truncated Spline Regression Using Maximum Likelihood Estimation for Human Development Disparities Laode Muhammad Saris; Henny Pramoedyo; Adji Achmad Rinaldo Fernandes
CAUCHY: Jurnal Matematika Murni dan Aplikasi Vol 10, No 1 (2025): CAUCHY: JURNAL MATEMATIKA MURNI DAN APLIKASI
Publisher : Mathematics Department, Maulana Malik Ibrahim State Islamic University of Malang

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.18860/cauchy.v10i1.31381

Abstract

A development of nonparametric truncated spline regression, Geographically Weighted Regression Spline Truncated (GWSTR) incorporates spatial effects in the modelling of nonlinear relationships between the response and predictor variables. This research utilizes the Maximum Likelihood Estimation (MLE) technique to estimate the parameters of the model. The first-order truncated spline with a single knot yielded a minimal Generalized cross-validation (GCV) value of 1. 729781, suggesting a high level of accuracy in the model.  Four weighting functions were evaluated: Gaussian Kernel, Exponential Kernel, Bi-Square Kernel, and Tri-Cube Kernel. Among these, the Bi-Square weighting function performed the best, achieving a coefficient of determination of 99.999%, demonstrating the model’s capability to explain nearly all data variability effectively. GWSTR proves to be a robust method for capturing complex nonlinear relationships while accounting for spatial variations, making it a valuable tool for spatial data analysis across various disciplines.
Development of Semiparametric Smoothing Spline Path Analysis on Cashless Society Muhammad Rafi Hasan Nurdin; Muhammad Ohid Ullah; Adji Achmad Rinaldo Fernandes; Eni Sumarminingsih; Solimun Solimun
CAUCHY: Jurnal Matematika Murni dan Aplikasi Vol 10, No 1 (2025): CAUCHY: JURNAL MATEMATIKA MURNI DAN APLIKASI
Publisher : Mathematics Department, Maulana Malik Ibrahim State Islamic University of Malang

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.18860/cauchy.v10i1.29846

Abstract

Path analysis requires assumptions to be met, particularly the linearity assumption, which can be tested using the Ramsey Regression Specification Error Test (RESET). Parametric path analysis is appropriate when all variable relationships are linear. For entirely non-linear relationships, a nonparametric model can be used, while a semiparametric model applies if there is a mix of linear and non-linear relationships. One nonparametric method is spline smoothing, which requires determining the spline polynomial order in estimating the nonparametric path function. Determining the spline polynomial order is challenging because there is no standard test for it. This study thus develops a modified Ramsey RESET to identify the optimal spline smoothing order. The development involves modifying the second regression equation with a nonparametric spline smoothing regression of orders 2 to 5. The modified Ramsey RESET algorithm is applied to cashless data, and the results are used to estimate a multi-group semiparametric smoothing spline function with a dummy variable approach. This estimation yields a goodness of fit of 94.14%, indicating that Product Quality and the Moderating Effect of Cashless Usage Frequency can explain Cashless User Satisfaction and Cashless User Loyalty by 94.14%, with the remaining 5.86% explained by variables outside the research model