Purpose – This study aims to develop and evaluate a phishing URL detection system using URL-based characteristics and multiple machine learning algorithms, while integrating Explainable Artificial Intelligence (XAI) to improve model transparency. Design/methods/approach – The study used 24,390 URLs consisting of 4,878 phishing URLs and 19,512 legitimate URLs collected from a public Kaggle dataset and manually collected sources. Thirteen URL-based features were extracted, including URL length, protocol usage, number of subdomains, domain length, IP address usage, and suspicious keywords. Six machine learning algorithms were compared: Random Forest, XGBoost, Decision Tree, Support Vector Machine, Logistic Regression, and Naïve Bayes. The dataset was evaluated using an 80:20 stratified train-test split, and model performance was measured using accuracy, precision, recall, F1-score, and confusion matrix. XAI was implemented using model-based feature importance and SHAP local explanations. Findings – XGBoost achieved the best overall performance with 97.89% accuracy, 93.09% precision, 96.62% recall, and 94.82% F1-score for the phishing class, followed closely by Random Forest. Feature importance analysis showed that protocol usage, suspicious keywords, and domain length were the most influential features. SHAP explanations further demonstrated that phishing predictions were influenced by the combined contribution of multiple URL characteristics. Research implications/limitations – The findings indicate that ensemble-based machine learning combined with explainable output can support transparent phishing URL screening. However, the study was limited to URL-based features, a fixed train-test split, and manually validated URLs without systematic verification from external threat-intelligence platforms. Originality/value – This study contributes by combining multi-algorithm phishing URL detection, URL-characteristic-based feature analysis, SHAP-based explanation, whitelist checking, and typosquatting indicators within an interpretable cybersecurity detection approach.
Copyrights © 2026