Claim Missing Document
Check
Articles

Found 1 Documents
Search

Comparison of Filter and SHAP Feature Selection for ECG-based Atrial Fibrillation Classification Novie Theresia Pasaribu; Elizabeth Fabiola Wijaya; Che Wei Lin
Jambura Journal of Electrical and Electronics Engineering Vol 8, No 2 (2026): Juli - Desember 2026
Publisher : Electrical Engineering Department Faculty of Engineering State University of Gorontalo

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.37905/jjeee.v8i2.39633

Abstract

Atrial Fibrillation (AF) is a type of arrhythmia whose prevalence continues to rise globally and can lead to serious complications such as stroke and heart attack. Early detection based on ECG signals is therefore crucial. This study compares three feature selection methods, namely Pearson Correlation (PC), Mutual Information (MI), and Shapley Additive Explanations (SHAP), for classifying AF from 2-lead ECG signals using XGBoost. The dataset used was the MIT-BIH Atrial Fibrillation Database, comprising 23 ECG recordings, yielding 34,312 10-second data segments. Preprocessing used Stationary Wavelet Transform (SWT) and Min-Max normalisation. A total of 50 features were extracted. Each method was tested on 4 feature subset sizes (5, 10, 15, and 20 features). The model was optimised using GridSearchCV with 5-fold Stratified Cross-Validation. Results showed that SHAP outperformed PC and MI in subsets of 10, 15, and 20 features. SHAP with 20 features achieved the highest performance (97.87% accuracy; F1 score 96.78%; ROC-AUC 0.9971), while the 15-feature SHAP offered the best performance–dimensionality trade-offs: 97.84% accuracy, 97.85% precision, 95.62% recall, 96.72% F1-score, and 0.9965 ROC-AUC, with a 70% dimensional reduction (0.03% below SHAP-20 on accuracy, with higher precision). SHAP's superiority over PC and MI was statistically significant (McNemar and DeLong tests, p 0.05) in subsets of 15 and 20 features; SHAP with 20 features achieved a ROC-AUC that did not differ significantly from the baseline of 50 features (DeLong test, p = 0.4997).