Journal of Defense Technology and Engineering
Vol. 2 No. 1 (2026): July, Journal of Defense Technology and Engineering

LightGBM based malware Classification for Cyber Defense Infrastructure Using the EMBER Portable Executable Feature Dataset

Hondor Saragih (Unhan)
Bartolomeu dos Reis (Dili Institute of Technology, Rua DIT Aimetilaran, Dili, Timor-Leste)
Vicente Soares (Dili Institute of Technology, Rua DIT Aimetilaran, Dili, Timor-Leste)



Article Info

Publish Date
31 Jul 2026

Abstract

Malware attacks against defense information systems continue to evolve in complexity, requiring automated multiclass classification methods capable of accurately distinguishing diverse malware families from high-dimensional Portable Executable (PE) features. This study presents a comprehensive evaluation of the LightGBM gradient boosting algorithm for nine-class malware classification using a stratified balanced subset of 405,000 samples from the EMBER benchmark dataset. The novelty of this work lies in providing a standardized comparative evaluation of LightGBM against XGBoost and Random Forest under identical data preprocessing, partitioning, and experimental conditions, enabling a fair assessment of the architectural advantages of each ensemble learning method. After removing zero-variance features, the dataset was represented by 2,350 static PE attributes and divided into training, validation, and testing subsets using a stratified 70/15/15 split. Experimental results demonstrate that the proposed LightGBM model achieved an overall accuracy of 98.12% and a macro F1-score of 98.12%, outperforming XGBoost and Random Forest under the same evaluation protocol. These findings indicate that LightGBM effectively captures complex feature interactions while maintaining computational efficiency for large-scale malware classification. The study contributes a reproducible benchmarking framework for multiclass malware classification and provides empirical evidence supporting the adoption of LightGBM as a practical ensemble learning approach for cyber defense applications. Future work will investigate the integration of dynamic behavioral features, explainable artificial intelligence techniques, and temporal malware datasets to improve model robustness against concept drift and adversarial attacks.

Copyrights © 2026