Publication Date
Spring 2026
Degree Type
Master's Project
Degree Name
Master of Science in Computer Science (MSCS)
Department
Computer Science
First Advisor
William B. Andreopoulos
Second Advisor
Amith Kamath Belman
Third Advisor
Pruthviraj Urankar
Keywords
large language models, prompt engineering, robustness, brittleness, retrieval-augmented generation, reproducibility
Abstract
Have you ever asked ChatGPT or Claude the same prompt twice and gotten different answers? Evaluating for brittleness based on one accuracy statistic takes the model’s skill to be the same thing as its average performance on one input, and this is incorrect. In this project, the gap between the two is measured using four families of tasks (MNLI, SST-2, ARC-Easy, and a Wikipedia retrieval task) where each individual prompt from the base family has been turned into variations (templates, paraphrases, personas, user profile, token edits, and document swap control). Across approximately 11,759 queries run through Gemini 2.5 Flash and Flash-Lite, brittleness proved to be a real phenomenon that varied per task: sentiment analysis was consistent, NLI benefited from a structured prompt, reasoning was highly dependent on prompt form, and retrieval was resilient except to paraphrasing.
Recommended Citation
Sanghvi, Shubh, "Evaluating Prompt Brittleness and Robustness in Large Language Models Across Classification, Reasoning, and Retrieval Tasks" (2026). Master's Projects. 1753.
DOI: https://doi.org/10.31979/etd.e6zn-8fbd
https://scholarworks.sjsu.edu/etd_projects/1753