Publication Date

Spring 2026

Degree Type

Master's Project

Degree Name

Master of Science in Computer Science (MSCS)

Department

Computer Science

First Advisor

Mark Stamp

Second Advisor

Fabio Di Troia

Third Advisor

Katerina Potika

Keywords

Image-based malware analysis, Explainable AI (XAI), Grad-CAM, Con- volutional neural networks (CNNs), Faithfulness metric, Stability metric

Abstract

Recent work has shown that binary-to-image malware representations can be used for effective machine learning–based malware classification, but performance varies significantly depending on the transformation technique used. A prior study evaluated several image transformations (grayscale, entropy images, and others) using traditional machine learning models and observed that the choice of transformation highly influences the classification accuracy. However, the explainability and interpretability of these models are unexplored for the most part. This project extends previous work in two phases. The first phase analyzes Grad-CAM heatmap behavior across eight image transformations and builds hybrid CNN+HOG+XGBoost models. The second phase introduces quantitative faithfulness and stability metrics for Grad- CAM evaluation and compares Grad-CAM with HiResCAM. A progressive feature combination approach extracts 256-dim CNN embeddings from models trained on both image types, and combining all 16 model embeddings into a 4096-dim feature vector achieves a test accuracy of 0.777 across 17 malware families, exceeding the prior benchmark of 0.750. An important finding is that the accuracy and explanation faithfulness are not directly correlated. The entropy_hcurve transformation offers the best overall balance, having high accuracy and a well balanced explainability.

Available for download on Saturday, May 22, 2027

Share

COinS