Publication Date

Spring 2026

Degree Type

Thesis

Degree Name

Master of Science (MS)

Department

Aviation and Technology

Advisor

Armin Moghadam; Mahima Agumbe Suresh; Riti Gour

Abstract

Enterprise authentication logs capture how users, hosts, and services interact over time, yet detecting cyber attacks in these logs remains difficult due to the rarity of malicious events, the absence of labeled training data, and the behavioral diversity across thousands of accounts. This thesis presents a three-stage pipeline for sequence-based anomaly detection and short-term attack-risk forecasting on the Los Alamos National Laboratory (LANL) Cyber1 dataset, which contains approximately 17 million authentication events spanning 58 days with 749 confirmed red team actions. In the first stage, a tuned Long Short-Term Memory (LSTM) network was trained exclusively on authentication sequences from benign days to perform next-token prediction over a 50-event context window, producing per-event surprisal scores without requiring attack labels. In the second stage, surprisal scores were normalized into per-user z-scores using baselines computed from confirmed benign days, then aggregated into three hourly features: maximum, mean, and 90th-percentile z-score. In the third stage, a LightGBM classifier with 100 fixed trees was trained on a nine-hour lookback window of these hourly features to forecast whether an attack would occur within the next one, three, or six hours. Evaluated under two temporal splits sharing identical test days, the primary model achieved PR-AUC values of 0.4325, 0.6390, and 0.9149 for horizons of one, three, and six hours, respectively, compared to random baselines of 0.1744, 0.3333, and 0.4815. Feature importance analysis revealed that at longer horizons, the oldest lags in the lookback window dominated, suggesting that the model captured signals consistent with multi-hour behavioral build-up patterns preceding attacks.

Available for download on Tuesday, January 26, 2027

Share

COinS