Publication Date
Spring 2026
Degree Type
Thesis
Degree Name
Master of Science (MS)
Department
Applied Data Science
Advisor
Mohammad Masum; Guannan Liu; Sayma Akhter
Abstract
Phishing through malicious URLs and deceptive emails remains a leading cybersecurity threat. While traditional machine learning classifiers achieve strong detection performance, their decisions are opaque and difficult for analysts to interpret or adapt to evolving attacks. This thesis investigates whether Large Language Models (LLMs) can provide competitive and explainable alternatives across both attack vectors. Phase I introduces the Feature-Enhanced Prompting for LLMs (FEP-LLM) framework, evaluating nine LLMs across ten prompting and feature transformation strategies for phishing URL classification. The best configuration achieves F1 = 94.09%, approaching the Random Forest baseline of 95.35%, with a key finding that feature augmentation benefit is model-capability-dependent, dramatically helping weaker models while degrading stronger ones. Phase II proposes a multi-agent framework for phishing email detection, comprising three role-specialized agents and a Meta-Judge synthesizer, with all components constrained to produce structured, schema-governed explanations. The system achieves Macro-F1 = 98.28% on a fixed 1,000-sample evaluation subset, outperforming a zero-shot single-model baseline by 6.3 percentage points while maintaining 99.45% phishing recall. Together, these studies establish that architectural design, through feature representation and evidence decomposition, determines LLM detection performance more than model capability alone, positioning LLMs as explainable, adaptive complements to traditional classifiers in cybersecurity workflows.
Recommended Citation
Yadav, Tanya, "From Feature-Enhanced Url Phishing Detection to Explainable Multi-Agent Email Phishing Detection Using Large Language Models" (2026). Master's Theses. 5806.
DOI: https://doi.org/10.31979/etd.azrh-tprx
https://scholarworks.sjsu.edu/etd_theses/5806