International Journal of Advances in Intelligent Informatics
Vol 12, No 3 (2026): August 2026

Adaptive hybrid ensemble-based DDoS detection using reinforcement learning-guided optimization and deep learning

Maha Ismail Raheem (College of Engineering, University of Information Technology and Communications, Al-Mansour, 3071, Baghdad)
Shouket Abdulrahman Ahmed (Department of Medical Instrumentation Techniques Engineering, Technical Engineering College, Al-Kitab University, Altun Kupri, Kirkuk, 36001)
Enas Faek Aziz (Department of Cybersecurity Engineering Technologies, Technical Engineering College for Computer and Artificial Intelligence/Kirkuk, Northern Technical University, 36001, Kirkuk)
Saad Ali Assi (Software Department, College of Computer Science and Information Technology, University of Kirkuk, Kirkuk)
Sinan Qahtan Salih (Technical College of Engineering, Al-Bayan University, Baghdad 10011)
Ahmed Dheyaa Radhi (College of Pharmacy, University of Al-Ameed, Karbala PO Box 198)
Hilal Adnan Fadhil (Department of Electrical and Computer Engineering, Sohar University, Sohar)
Taha Almulaisi (Renewable Energy Research Unit, Polytechnic College Hawija, Northern Technical University, Hawija, 36007)



Article Info

Publish Date
31 Aug 2026

Abstract

Distributed Denial-of-Service (DDoS) attacks remain among the most disruptive network threats, and detectors that generalize across attack families with low false-alarm rates are still an open problem. Propose an adaptive hybrid ensemble that unifies two gradient-boosting learners (Random Forest and Gradient Boosting) with three deep neural base learners (DNN, CNN-1D, and LSTM) under a weighted soft-voting rule whose weights are produced by a Reinforcement Learning (RL) policy. The RL agent treats the ensemble-weight simplex as its action space, observes a state vector built from validation-set diagnostic statistics, and is trained by REINFORCE-with-baseline to maximize a reward equal to validation F1 minus a small calibration penalty. The framework is formalized as a Markov decision process with one stochastic step per training episode, which decouples ensemble-weight learning from the (non-differentiable) outer F1 objective. On a 10,000-sample, 25-feature, five-class benchmark with 7% label noise, the proposed system reaches weighted F1 = 0.846, accuracy = 84.7%, MCC = 0.781, AUC = 0.952, and ECE = 0.039. Friedman and Nemenyi post-hoc tests over 50 CV folds confirm the RL-guided ensemble is significantly better than every individual base learner and uniform voting at α = 0.05 (Cohen's d = 0.96). An ablation isolates the RL policy and gradient boosting as the main drivers; a label-noise robustness study shows graceful degradation up to 20%; a head-to-head comparison against the Bonobo Optimizer (BO), GA, PSO, GWO, and WOA shows the best F1/wallclock trade-off.

Copyrights © 2026






Journal Info

Abbrev

IJAIN

Publisher

Subject

Computer Science & IT

Description

International journal of advances in intelligent informatics (IJAIN) e-ISSN: 2442-6571 is a peer reviewed open-access journal published three times a year in English-language, provides scientists and engineers throughout the world for the exchange and dissemination of theoretical and ...