Phishing URL detection using machine learning often achieves high performance, but feature relevance may vary across datasets and collection contexts. This study analyzes the characteristics of an Indonesian phishing URL dataset and evaluates a compact detection model that prioritizes phishing recall. The dataset contains 2,054 balanced phishing and legitimate URLs represented by 70 lexical, structural, domain/WHOIS, infrastructure, and reputation features. Interquartile Range (IQR) was used to identify statistically extreme values, while Pearson correlation and Mutual Information were used to examine feature-target associations. LASSO+ was applied to reduce redundant features, and Random Forest with uncertainty-weighted bootstrap sampling was evaluated using stratified 10-fold cross-validation. The results show that 92.84% of the URLs contain at least one statistically extreme feature value. Response time, IP information, domain reputation, and hyperlink-related features show prominent associations with the target. LASSO+ reduces the feature set by approximately 36.3% without reducing Random Forest performance. The complete model achieves the highest recall of 0.9942 ± 0.0050 and the lowest false-negative count of six. These findings show that feature relevance needs to be validated for the target dataset and that a compact feature representation can support recall-oriented phishing detection.
Copyrights © 2026