Web-Based Phishing Detection in the Era of Explainable and Generative AI: A Critical Narrative Review of Machine Learning, Deep Learning, Explainable AI and Large Language Models
C. Uzoaru Godson
Department of Maths and Computer Science, Clifford University, Owerrinta, Abia State, Nigeria.
Nwasuka Stanley
Department of Maths and Computer Science, Clifford University, Owerrinta, Abia State, Nigeria.
G. U. Nwamuruamu
Department of Maths and Computer Science, Clifford University, Owerrinta, Abia State, Nigeria.
Enemanna Princess-Precious *
Department of Maths and Computer Science, Clifford University, Owerrinta, Abia State, Nigeria.
C. Johnson Francis
Department of Maths and Computer Science, Clifford University, Owerrinta, Abia State, Nigeria.
A. Keke Ndu Levis
Department of Maths and Computer Science, Clifford University, Owerrinta, Abia State, Nigeria.
Uwakonye Ogbuja Justine
Department of Maths and Computer Science, Clifford University, Owerrinta, Abia State, Nigeria.
Onyenkwere Chinemerem Divine
Department of Maths and Computer Science, Clifford University, Owerrinta, Abia State, Nigeria.
*Author to whom correspondence should be addressed.
Abstract
Web-based phishing detection has progressed from manually specified URL and webpage heuristics to machine learning, deep representation learning, explainable artificial intelligence, and, most recently, large language and multimodal models. This technical expansion has produced impressive benchmark results but has also made comparison increasingly difficult because models are evaluated on different datasets, temporal windows, feature sets, class balances, and operational assumptions. This critical narrative review evaluates the evidence for machine learning, deep learning, explainable artificial intelligence and large language model approaches to detecting phishing URLs and webpages, with emphasis on generalisability, adversarial robustness, interpretability, computational cost and deployment relevance. Literature published principally from 2010 to 9 June 2026 was identified through accessible scholarly indexes, digital-library records, DOI metadata, citation searching and authoritative proceedings. The evidence indicates that conventional machine-learning pipelines remain competitive when lexical, host, hyperlink and structural features are well engineered, particularly where latency and transparency matter. Deep learning reduces dependence on handcrafted features and can exploit sequential, visual and multimodal representations, but reported gains are often conditional on dataset construction and may weaken under temporal shift or adversarial manipulation. Explainable artificial intelligence improves model inspection and can reveal influential phishing cues, yet post-hoc explanations do not by themselves establish causal validity, stability or security robustness. Large language and multimodal models add semantic reasoning, brand identification and natural-language explanation, but the evidence base remains recent and operationally constrained by cost, latency, privacy, reproducibility and prompt- or content-mediated attack surfaces. Across paradigms, the dominant unresolved problem is not benchmark accuracy but reliable performance under changing, adversarial web conditions. The review therefore argues for temporally separated, cross-source and adversarial evaluation, calibrated uncertainty, explanation-quality testing and layered architectures in which lightweight detectors, visual or content models and language-model reasoning are invoked according to risk and resource constraints.
Keywords: Anti-phishing, malicious URL classification, webpage security, adversarial robustness, model interpretability, multimodal learning, generative artificial intelligence, cybersecurity