Author

Publication Date

Spring 2026

Degree Type

Master's Project

Degree Name

Master of Science in Computer Science (MSCS)

Department

Computer Science

First Advisor

William B. Andreopoulos

Second Advisor

Amith Kamath Belman

Third Advisor

Pruthviraj Urankar

Keywords

large language models, prompt engineering, robustness, brittleness, retrieval-augmented generation, reproducibility

Abstract

Have you ever asked ChatGPT or Claude the same prompt twice and gotten different answers? Evaluating for brittleness based on one accuracy statistic takes the model’s skill to be the same thing as its average performance on one input, and this is incorrect. In this project, the gap between the two is measured using four families of tasks (MNLI, SST-2, ARC-Easy, and a Wikipedia retrieval task) where each individual prompt from the base family has been turned into variations (templates, paraphrases, personas, user profile, token edits, and document swap control). Across approximately 11,759 queries run through Gemini 2.5 Flash and Flash-Lite, brittleness proved to be a real phenomenon that varied per task: sentiment analysis was consistent, NLI benefited from a structured prompt, reasoning was highly dependent on prompt form, and retrieval was resilient except to paraphrasing.

Available for download on Saturday, May 22, 2027

Share

COinS