Cybersecurity Analysis: Explainable Machine Learning for Phishing URL Detection

Authors

  • Aishwarya Sawant Chantilly High School, Chantilly, VA
  • Nikil Pillai Thomas Jefferson High School for Science and Technology, Alexandria, VA
  • Kamaljeet Sanghera Department of Information Sciences and Technology, George Mason University, Fairfax, VA

DOI:

https://doi.org/10.13021/jssr2026.5695

Abstract

Phishing websites present one of the biggest security concerns due to their nature of imitating legitimate websites and theft of personal information through such actions. In this project, we explore machine learning models' ability to detect phishing websites based on their characteristics, such as URLs and webpages. The analysis was carried out using the PhiUSIIL Phishing URL Dataset and included 235,795 URLs with 54 features. Two machine learning models, Logistic Regression and Random Forest, were trained using an 80/20 split of the data. Performance of the models was measured by accuracy, precision, recall, F1 score, ROC curves, and a confusion matrix. Both models achieved excellent results when applied to testing data. Logistic Regression achieved an accuracy of 99.987%, and Random Forest achieved an accuracy of 100%. Additionally, feature importance was estimated for the Random Forest model. Therefore, the results obtained in this project show the effectiveness of machine learning in detecting phishing websites in the dataset used. Still, such results need to be treated cautiously since good performance on a benchmark dataset may not reflect performance on real-world phishing websites.

Published

2026-09-24

Issue

Section

College of Engineering and Computing: Department of Information Sciences and Technology