Farizal Herry Saputra
Pamulang University

Published : 1 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 1 Documents
Search

Analysis and Development of Transformer (DistilBERT)-LSTM and Reinforcement Learning Models for Adaptive Phishing Email Detection Farizal Herry Saputra; Kahfi Heryandi Suradiraja; Abu Khalid Rivai
SISTEMASI Vol 15, No 7 (2026): Sistemasi: Jurnal Sistem Informasi
Publisher : Universitas Islam Indragiri

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.32520/stmsi.v15i7.6432

Abstract

Phishing detection faces two major challenges: performance degradation caused by domain shift and a heavy reliance on costly labeled data. This study proposes an adaptive phishing email detection model that integrates a hybrid DistilBERT–LSTM architecture with a Proximal Policy Optimization (PPO)-based Reinforcement Learning agent. The proposed methodology employs a multi-stage transfer learning framework using three datasets: Enron as the source domain, Phishing_Validation for supervised domain adaptation, and CEAS_08 to simulate an unlabeled data stream through pseudo-labeling. Experimental results demonstrate excellent performance on the source-domain dataset (Enron), achieving an F1-score of 0.9935. However, the model's performance declined on the Phishing_Validation dataset (F1-score = 0.9067), confirming the impact of domain shift. By incorporating the PPO agent, the proposed model autonomously recovered its performance on the CEAS_08 dataset, achieving an F1-score of 0.9516, an accuracy of 0.9468, and a ROC–AUC of 0.9915. The stability of the adaptation process was validated by the convergence of the Kullback–Leibler (KL) divergence to 0.000173, although a minor overconfidence of approximately 5% was observed between the model's confidence estimates and the ground-truth labels. These findings demonstrate the effectiveness of PPO in mitigating domain shift within an unsupervised adaptation environment. Future research should focus on improving model calibration and exploring multimodal feature integration to further strengthen cybersecurity defenses.