Claim Missing Document
Check
Articles

Improving Internet of Things Cyber Attack Detection with Information Gain and Decision Tree Alvin Mufidha Ahmad; Fauzi Adi Rafrastara
Journal of Innovation and Technology Polbeng Series on Informatics (INOVTEK Polbeng - Seri Informatika) Vol. 10 No. 3 (2025): November
Publisher : P3M Politeknik Negeri Bengkalis

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.35314/p2c33t87

Abstract

The rapid proliferation of Internet of Things (IoT) devices within digital ecosystems has enhanced efficiency and availability but has also expanded the attack surface for cyber threats. This study aims to improve intrusion detection accuracy in IoT environments by addressing two key challenges: class imbalance and high feature dimensionality. Random Undersampling (RUS) is employed to mitigate data imbalance in the CIC IoT 2023 dataset, while feature selection is performed using the filter-based Information Gain method. A decision tree classifier is implemented and validated using k-fold cross-validation to ensure result reliability. Experimental results demonstrate that the proposed approach achieves an accuracy of 88.7%, outperforming a wrapper-based method, which attained 87.3%. These findings confirm that an appropriately designed filter-based feature selection strategy can effectively enhance the performance of intrusion detection systems for IoT security.
A Comparative Analysis of P-Value and Mutual Information Feature Selection Methods for Random Forest-Based Phishing Detection Fahmi Bahtiar Adi Nugroho; Wildanil Ghozi; Fauzi Adi Rafrastara
Jurnal Nasional Teknologi dan Sistem Informasi Vol 11 No 3 (2025): Desember 2025
Publisher : Departemen Sistem Informasi, Fakultas Teknologi Informasi, Universitas Andalas

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.25077/TEKNOSI.v11i3.2025.377-386

Abstract

The application of ANOVA's P-Value-based feature selection method, namely the F-test, in phishing detection with the Random Forest algorithm indicates that a configuration of 25 features yields the quickest inference time, rendering it appropriate for scenarios demanding great computational efficiency and responsiveness. However, if the user's primary priority is to achieve the highest level of detection accuracy, the 29-feature configuration is more feasible because it exhibits higher accuracy performance and better prediction stability. Consequently, there is no definitive trade-off between 25 or 29 features, there exists a selection of solutions that can be tailored to the application's requirements. This methodology enables users to achieve an optimal equilibrium between superior performance and minimal inference time in a phishing detection system, contingent upon the implementation context and operational priorities. This study successfully shows that a simple statistical approach such as P-Value is not only competitive but also provides superior results compared to more complex methods, offering a practical and efficient solution for real-world implementation.