Publication Date

Spring 2026

Degree Type

Master's Project

Degree Name

Master of Science in Computer Science (MSCS)

Department

Computer Science

First Advisor

Navrati Saxena

Second Advisor

Genya Ishigaki

Third Advisor

Kalindi Parekh

Keywords

clinical reasoning, large language models, retrieval augmented generation, multi-agent systems, hallucination reduction

Abstract

Recent advances in large language models have enabled strong performance on medical QA benchmarks, yet their clinical reliability remains limited due to unsup- ported reasoning, incomplete differential diagnoses, and single pass inference. We propose a retrieval augmented multi agent framework that improves reasoning through three stages: evidence conditioned hypothesis generation, adversarial critique to chal- lenge hypotheses, and independent verification. Each component is grounded using semantic retrieval from PubMed, MedQA explanations, and MIMIC III records. The system is evaluated on MedQA USMLE and MIMIC derived prompts against both a general purpose LLM and a fine tuned medical model. Results show improved supported sentence rate, attribution precision and recall, and overall evidence ground- edness, alongside reductions in unsupported claims, hallucinations, Brier score, and expected calibration error, while maintaining diagnostic accuracy. This demonstrates that adversarial and verification based inference enhances evidence alignment and uncertainty calibration without requiring additional supervised training.

Available for download on Saturday, May 22, 2027

Share

COinS