Claim Missing Document
Check
Articles

Found 1 Documents
Search

A Comparative Evaluation of Model-Based and SHAP-Based Feature Analysis for Robust Health Data Classification Indra Waspada; Satriawan Rasyid Purnama; Alwey Hakim; Alfonso Clement Sutantio
Scientific Journal of Informatics Vol. 13 No. 2: May 2026
Publisher : Universitas Negeri Semarang

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.15294/sji.v13i2.42406

Abstract

Purpose: Feature selection is one critical element of healthcare data classification, which directly affects predictive performance, model robustness, and interpretability. Nevertheless, traditional model-based feature importance methods are unstable in robustness and provide random or misleading results on high-dimensional and heterogeneous healthcare data. The purpose of this paper was to assess model-based and SHapley Additive exPlanations (SHAP) - based feature analysis on multimodal healthcare data classification in a comparative manner. Methods: This study employed a quantitative comparative experimental design using real Electronic Medical Records (EMR) data captured from primary care clinics. The dataset comprises 2,158 patient records, with numerical and textual features in a multimodal feature space. The analytical pipeline included data acquisition, preprocessing, multimodal feature integration, and model development using six supervised learning algorithms from ensemble-based and margin-based categories. The model-based feature importance was used with the SHAP-based feature importance. We examined the robustness of this method across different scenarios through systematic feature ablation using baseline, strong, and weak features. Result: Experimental results demonstrate that model-based feature importance exhibits unpredictable behavior when features are removed and is sometimes counterintuitive. On the other hand, feature importance based on SHAP is consistent and linearly proportional across the models under consideration. The highest macro F1-score was achieved by Extra Trees with SHAP-based strong features (0.824), exceeding the baseline (0.811). In robustness testing, SHAP-based weak-feature removal reduced the LightGBM F1-score from 0.764 to 0.575, suggesting a clearer distinction between informative and non-informative features. Overall, SHAP-based feature selection provided a more reliable and interpretable framework for multimodal healthcare classification. Novelty: Our experiment results demonstrate empirically that the SHAP-based feature importance is more robust and reliable than conventional model-based approaches when it comes to extracting features from multimodal medical records. This work shows that SHAP is more than a post hoc explanation by presenting it as an interpretable feature selection criterion guiding feature relevance analysis in healthcare machine learning.