Publication Date

Spring 2026

Degree Type

Master's Project

Degree Name

Master of Science in Computer Science (MSCS)

Department

Computer Science

First Advisor

Navrati Saxena

Second Advisor

Robert Chun

Third Advisor

Abhishek Roy

Keywords

Retrieval-Augmented Generation, Educational AI, Curriculum- Aware Evaluation, Large Language Models, AI Tutoring Systems, Evaluation Metrics

Abstract

Retrieval-Augmented Generation (RAG) systems are being increasingly used as tutoring assistants. Current evaluation frameworks mainly focus on factual correctness and retrieval grounding, but they do not evaluate if the response is appropriate for a learner’s given position in the curriculum. To address this gap, this paper introduces a temporal evaluation framework that measures pedagogical scope compliance in RAG-based tutoring systems. The approach consists of three components. An automated Knowledge Component extraction pipeline that builds the curriculum concept timeline from lecture transcripts with inter-judge agreement �� = 0.84, a diagnostic benchmark of 187 LLM-generated and validated questions designed to evaluate scope compliance, and a metric suite including Scope Violation Rate, answer utility, and question-level success to capture the balance between scope compliance and answer helpfulness. This paper evaluates five retrieval and prompting conditions across two generation models (Claude Sonnet 4.5 and GPT-4o) on the MIT 6.034 Artificial Intelligence course. Across both models, hard retrieval filtering consistently reduces scope violations by 69% relative to the unrestricted baseline while maintaining answer quality, whereas instruction-based suppression methods are less consistent across models. Overall, the framework produces a reproducible methodology for evaluating whether RAG tutoring systems respect curricular scope, addressing a limitation that standard correctness-based metrics don’t capture.

Available for download on Wednesday, May 19, 2027

Share

COinS