Fitriana AR
Department of Statistics, Universitas Syiah Kuala, Banda Aceh 23111, Indonesia

Published : 1 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 1 Documents
Search

Ensemble Variable Importance: Combining Random Forest, Neural Network, and Support Vector Machine via Genetic Algorithm (Case Study: Student Productivity) Asep Rusyana; Marzuki Marzuki; Siti Rusdiana; Fitriana AR; Nurhasanah Nurhasanah; Nany Salwa; Mahmudi Mahmudi
Infolitika Journal of Data Science Vol. 4 No. 1 (2026): May 2026
Publisher : Heca Sentra Analitika

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.60084/ijds.v4i1.424

Abstract

This study proposes and evaluates an ensemble variable‑importance framework that integrates permutation‑based importance scores from three distinct supervised learning algorithms: Random Forest, Neural Network, and Support Vector Machine, using a genetic‑algorithm optimizer. The approach addresses the well‑known problem that algorithm‑specific importance diagnostics can yield divergent feature rankings, complicating substantive interpretation and downstream decision‑making. Using a large publicly available student‑productivity dataset (N = 20,000), predictors describing study behavior, digital‑media use, lifestyle, and academic indicators were normalized with Min–Max scaling, and permutation variable importance (PVI) was estimated repeatedly within each model to obtain stable mean PVI values and standard errors. A genetic algorithm was then employed to search the space of ensemble weightings (rank‑aggregation solutions) that maximize a chosen fitness criterion—Spearman rank concordance with out‑of‑sample predictive relevance—thereby producing a consensus ranking of predictors. Empirical results indicate rapid GA convergence (fitness ≈ 0.82 within 20–30 generations) and strong cross‑model agreement for a small core of predictors: study hours (X3) and focus score (X15) consistently emerged as the most salient features across individual models and in the ensemble ranking. A secondary set of variables (e.g., sleep hours, phone usage, attendance, and stress level) displayed moderate importance, while several features exhibited model‑dependent variability in ranks. The ensemble procedure thereby yields stable, model‑agnostic importance estimates that enhance interpretability and reduce dependence on any single algorithm’s idiosyncrasies. We discuss implications for educational analytics and recommend external validation, targeted feature engineering, and sensitivity analyses (alternate scalings and GA settings) to assess robustness and to support reliable, actionable inferences from machine‑learning models in applied settings.