Rachana Sandeep Potpelwar, J. M. Waghmare and U. V. Kulkarni
Adv. Artif. Intell. Mach. Learn., 6 (4):6108-6121
1. Rachana Sandeep Potpelwar: Information Technology Department Shri Guru Gobind Singhji Institute of Engineering and Technology, Vishnupuri, India
2. J. M. Waghmare: Department of Computer Science and Engineering Shri Guru Gobind Singhji Institute of Engineering and Technology, Vishnupuri, India
3. U. V. Kulkarni: Department of Computer Science and Engineering Shri Guru Gobind Singhji Institute of Engineering and Technology, Vishnupuri, India.
DOI: 10.54364/AAIML.2026.64338
Article History: Received on: 23-Feb-26, Accepted on: 21-Aug-26, Published on: 28-Aug-26
Corresponding Author: Rachana Sandeep Potpelwar
Email: rspotpelwar@sggs.ac.in
Citation: Rachana Sandeep Potpelwar, et al. OFSPUD: Optimizing Feature Selection for Phishing URL Detection through Statistical Tests and Ensemble Methods. Advances in Artificial Intelligence and Machine Learning. 2026Íľ6(4):338. https://dx.doi.org/10.54364/AAIML.2026.64338
The surge in online users utilizing cloud-based platforms—particularly for financial
and retail services—is fueled by the convenience and flexibility these services provide.
To effectively counter evolving URL phishing threats, machine learning classifiers
should analyze all components of URLs, including domain structures, path patterns,
and query parameters, to enhance threat detection. The proposed framework, OFSPUD
(Optimizing Feature Selection for Phishing URL Detection through Statistical Tests
and Ensemble Methods), was evaluated using a large, imbalanced dataset containing
11,055 benign and phishing URLs with over 31 features. Using ANOVA (Analysis of
Variance), 25 of the 31 features were selected, significantly improving the method’s
performance and correctness. The RBWSV (Rank-Based Weighted Soft Voting) en
semble method outperformed prominent machine learning techniques, achieving test
accuracy of 97.47% and F1 and validation accuracies of 99% after applying ANOVA
based feature selection. Thus, the proposed method, combining ANOVA-based feature
selection with the RBWSV ensemble approach, achieved near-perfect accuracy while
improving efficiency.