Publication Date

Spring 2026

Degree Type

Master's Project

Degree Name

Master of Science in Computer Science (MSCS)

Department

Computer Science

First Advisor

William B. Andreopoulos

Second Advisor

Faranak Abri

Third Advisor

Navrati Saxena

Keywords

context engineering; local RAG; perturbation robustness; decision- gate; perplexity; LLM as a judge

Abstract

Local Retrieval-Augmented Generation can change behavior under tiny prompt edits or context reorderings, yet practical testing often uses only one query formulation. This report presents a local context-engineering framework for perturbation robustness, reproducibility, and budget-aware evaluation. The system logs each run as a structured capsule, computes lexical, retrieval, fluency, and semantic metrics, and uses a calibrated

decision-gate to skip, reduce, or execute the perturbation suite. The augmented capsule- derived dataset has 3,570 perturbation rows. The retraining was done on 2,619 rows from the augmented data with enriched observed-break labels; 446 of those rows had BLEU scores, 1,668 had perplexity values, and 1,668 had semantic-judge output values. The trained gate achieved 0.915 ROC-AUC, 0.891 PR-AUC, 0.913 accuracy, 0.863 F1 score, and 76.5% expected cost savings on the held-out augmented test split.

Available for download on Saturday, May 22, 2027

Share

COinS