Artificial Intelligence for Phishing and Social-Engineering Detection: A Machine Learning Evaluation on a Benchmark URL Dataset

Authors

  • Muhammad Azeem Afzal Department of Cyber Security, NASTP Institute of Information Technology, Lahore, Pakistan Author
  • Usama Ahmad Mughal Department of Cyber Security, NASTP Institute of Information Technology, Lahore, Pakistan Author
  • Malik Hammad Hussain Department of Software Engineering, NASTP Institute of Information Technology, Lahore, Pakistan Author
  • Muhammad Daniyal Qadri Department of Artificial Intelligence, NASTP Institute of Information Technology, Lahore, Pakistan Author
  • Muhammad Daniyal Baig Department of Artificial Intelligence, NASTP Institute of Information Technology, Lahore, Pakistan Author
  • Muhammad Ehsan Qadeer Department of Mechanical Engineering, University of Lahore, Lahore, Pakistan Author

DOI:

https://doi.org/10.57041/c8cf4x04

Keywords:

Phishing detection, machine learning, social engineering, large language models, adversarial robustness, ensemble classifiers

Abstract

Phishing remains the leading initial-access vector in modern cyberattacks, and the emergence of large language models (LLMs) has made social-engineering content easier to produce and harder to distinguish from legitimate communication. This paper reviews recent (2023-2026) research on artificial-intelligence-based phishing detection and reports an original empirical evaluation on a public benchmark dataset of 88,647 labelled URLs described by 111 engineered features spanning URL, domain, directory/file, query-parameter, and network/WHOIS-derived attributes. Five supervised learning models logistic regression, linear support vector machine, random forest, gradient boosting, and a multilayer perceptron were trained and evaluated using stratified five-fold cross-validation. The random forest classifier achieved the strongest performance (accuracy = 0.970, F1 = 0.957, ROC-AUC = 0.995), consistent with the ensemble-dominance trend reported in the recent literature. A follow-up robustness simulation shows that although the model is highly resistant to random feature-level noise, this does not imply resistance to the targeted, semantically-aware evasion strategies enabled by generative AI, which recent studies show can bypass commercial filters. The paper concludes with a discussion of the dual-use nature of AI in this domain and directions for more adversarial-robust, content-aware detection systems.

Downloads

Published

2026-09-21

How to Cite

Artificial Intelligence for Phishing and Social-Engineering Detection: A Machine Learning Evaluation on a Benchmark URL Dataset. (2026). International Journal of Emerging Engineering and Technology, 5(1-1), 19-25. https://doi.org/10.57041/c8cf4x04

Most read articles by the same author(s)

Similar Articles

1-10 of 53

You may also start an advanced similarity search for this article.