HRIT: A Human-Readable Framework for Phishing URL Detection Using Large Language Models

Publication Date

1-1-2026

Document Type

Conference Proceeding

Publication Title

2026 IEEE 5th International Conference on AI in Cybersecurity Icaic 2026

DOI

10.1109/ICAIC67076.2026.11395727

Abstract

Phishing URL detection remains a persistent cybersecurity challenge due to rapidly evolving adversarial tactics and the inherent limitations of rule-based and traditional machine learning approaches. Although Large Language Models offer promising capabilities in contextual reasoning and natural-language explanation, their effective use for URL-based threat detection is constrained by a semantic mismatch between numeric security features and language-model reasoning. This paper introduces HRIT (Human-Readable Indicator Transformation), a lightweight and model-agnostic framework that transforms structured URL security features into semantically meaningful, human-readable indicators optimized for LLM inference. HRIT enables LLMs to reason about phishing intent using interpretable descriptors derived from statistically validated, dataset-driven, and industry-informed thresholds, without requiring model fine-tuning or retraining. We evaluate HRIT on a balanced phishing URL dataset containing 11,430 samples and compare its performance against a strong Random Forest baseline and multiple LLMs under zero-shot prompting. Experimental results show that HRIT improves detection sensitivity, improving recall and F1-score while maintaining low inference cost and latency. These findings demonstrate that effective phishing URL detection can be achieved through semantic feature transformation rather than complex prompting or model adaptation.

Keywords

Cybersecurity, Feature Transformation, Human-Readable Indicators, Large Language Models, Phishing URL Detection, Prompt Engineering

Department

Applied Data Science

Share

COinS