Journal of Applied Data Sciences
Vol 7, No 3: September 2026

Analysis of Machine Learning Models Based on Predictive Performance, Energy Consumption, and Carbon Emissions

Willy Permana Putra (Politeknik Negeri Indramayu)
Eko Marpanaji (Unknown)
Septafiansyah Dwi Putra (Unknown)



Article Info

Publish Date
10 Aug 2026

Abstract

This study aims to evaluate intrusion detection models by jointly considering predictive performance and computational sustainability. The main problem addressed is that many intrusion detection studies emphasize classification accuracy while providing limited evidence about execution time, energy use, and carbon dioxide-equivalent emissions, even though these factors affect repeated training and practical deployment. The contribution of this work is a comparative assessment of four supervised learning models, Random Forest, Histogram-based Gradient Boosting, Support Vector Machine, and Extreme Gradient Boosting, under a unified experimental workflow. The methodology uses the Wednesday working-hours subset of the Canadian Institute for Cybersecurity intrusion detection dataset released in 2017, which contains benign traffic and several denial-of-service attack classes. The procedure includes dataset selection, data cleaning, preprocessing, stratified training and testing, model fitting, predictive evaluation, sustainability measurement, and comparative interpretation. The evaluation is supported by one workflow figure, tables describing predictive results and sustainability measurements, and comparative visualizations of classification and resource-efficiency outcomes. The results show that Extreme Gradient Boosting achieved the strongest overall classification performance, with an accuracy of 0.9994, macro recall of 0.9966, macro F1-score of 0.9956, macro precision of 0.9946, and one-versus-rest area under the curve of 0.9999, while requiring 0.002559 kilowatt-hours of energy and 11.83 seconds of execution time. Random Forest produced highly comparable predictive results with similarly low resource consumption. Histogram-based Gradient Boosting was the most efficient model in terms of time, energy use, and emissions, but its macro-level performance was substantially lower. Support Vector Machine produced acceptable predictive results but required substantially higher computational resources. These findings imply that sustainable intrusion detection should select models through a balanced evaluation of detection capability and computational cost rather than accuracy alone.

Copyrights © 2026






Journal Info

Abbrev

JADS

Publisher

Subject

Computer Science & IT Control & Systems Engineering Decision Sciences, Operations Research & Management

Description

One of the current hot topics in science is data: how can datasets be used in scientific and scholarly research in a more reliable, citable and accountable way? Data is of paramount importance to scientific progress, yet most research data remains private. Enhancing the transparency of the processes ...