Publication Date

10-9-2026

Document Type

Article

Publication Title

Knowledge Based Systems

Volume

351

DOI

10.1016/j.knosys.2026.116854

Abstract

Large Language Models (LLMs) possess strong general-purpose reasoning abilities, yet their integration into static malware analysis pipelines remains limited due to the representational gap between high-dimensional numerical features and the natural-language input format required for inference. This study introduces PE2Prompt, a Large Language Model-native framework for interpretable and industrially deployable static malware analysis. PE2Prompt formalizes a modular process that bridges numeric PE features and LLM reasoning through three key components: (1) LLM-guided feature curation directly from raw EMBER JSON fields to identify semantically relevant attributes; (2) numerical feature summarization, compressing byte-histogram and entropy distributions into interpretable descriptors; and (3) structured prompt construction, transforming curated PE semantics into language-based reasoning templates. We evaluate zero-shot and few-shot prompting across GPT-4o-mini, Claude 3.7 Sonnet, and DeepSeek-Chat, analyzing classification accuracy, recall asymmetry, latency, and inference cost under varying feature-selection thresholds. Claude achieves the most balanced performance (≈73% accuracy) while DeepSeek-Chat demonstrates improved stability under structured prompting (≈59%), and GPT-4o-mini achieves below-random accuracy (≈49%) with a pronounced benign bias. Beyond empirical performance, PE2Prompt establishes a reproducible methodology for integrating LLM reasoning within static malware workflows, defining a new feature-to-language reasoning paradigm that advances explainable, adaptable, and zero-retraining malware analysis for industrial cybersecurity environments. To facilitate reproducibility, transparency, and future benchmarking, all source code, prompting pipelines, and experimental scripts are publicly available through an open-source repository.

Keywords

EMBER dataset, Large language models, Prompt engineering, Static malware detection, Zero-shot classification

Creative Commons License

Creative Commons License
This work is licensed under a Creative Commons Attribution-Noncommercial 4.0 License

Department

Applied Data Science

Share

COinS