Text based social engineering attacks are a growing cyber threat that is difficult to detect by conventional intrusion detection systems, especially in previously unobserved or zero-day variants. This study proposes a Natural Language Processing Open-Set Intrusion Detection System (NLP-OSIDS) framework that integrates Term Frequency-Inverse Document Frequency (TF-IDF) trigram (1.3-gram) feature representation with an Open-Set Multilayer Perceptron architecture based on energy based scoring to detect zero-day social engineering attacks without requiring training examples from that class. Experiments were conducted on the public dataset phishing_email.csv with 82,486 combined samples from Enron, SpamAssassin, Nazario, Ling, CEAS, and Nigerian Fraud datasets with strict zero-day partitioning following open-set recognition evaluation standards. The results show that NLP-OSIDS achieved an AUROC of 0.7808, surpassing all closed-set baselines (AUROC = 0.500) with the lowest False Positive Rate of 0.0088, while the Zero-Day Detection Rate (ZD-DR) of 0.077 indicates the need for adaptive threshold optimization as a direction for further research.
Copyrights © 2026