Publication Date
Spring 2026
Degree Type
Master's Project
Degree Name
Master of Science in Computer Science (MSCS)
Department
Computer Science
First Advisor
Amith Kamath Belman
Second Advisor
Wendy Lee
Third Advisor
Philip Heller
Keywords
Item Response Theory, LLM Evaluation, Multi-Judge Consensus, Adaptive Tests, Computerized Adaptive Tests, LLM-as-Judge
Abstract
Item Response Theory (IRT) has been used for evaluating LLMs using adaptive test-taking. Balkır et al. showed that a 1-parameter logistic (Rasch) model combined with computerized adaptive testing allows obtaining a cost-efficient LLM ability estimation on open-ended tasks. A clear disadvantage of their method is using a single-score metric, which becomes a source of error in measurement noise. In the present study, we use the multi-judge consensus scoring approach: a set of calibrated LLM evaluators providing continuous IRT responses in the context of adaptive tests. For a 1PL fixed-discrimination model, variance in scoring is the only parameter available besides item selection to improve the precision of ability estimation. Multi-judge consensus scoring decreases this score variance by averaging across independent scoring sources, producing more accurate ability estimates from fewer items. Our experiments with two benchmarks, TatQA (financial tabular reasoning) and MT-Bench-101 (multi-turn dialogue), using three local LLMs (Qwen2-Math 1.5B, Llama 3.2 3B Instruct, and Qwen2.5 1.5B Instruct) and two consensus judges (GPT-4 and DeepSeek), show high reproducibility in agent ranking on both datasets. With MT-Bench-101, we obtained non-overlapping 95% confidence intervals with ∆ θ≤0.08 logits for all agents in repeated runs. Meanwhile, on TatQA, Qwen2-Math exhibits higher variance but still has consistent relative agent ranking. When the score variance of our estimations was adjusted, all settings converged in 8 to 16 items. However, using one judge makes the ability distribution dependent on the judge itself: GPT-4 inflates ability by 0.35 to 2.5 logits depending on the agent. Using multi-judge consensus makes the measurement less dependent on the behavior of one judge, reducing standard errors by 18% to 24% compared to using a single judge.
Recommended Citation
Podey, Omkar, "Multi-Judge Consensus Scoring for IRT-Based Adaptive LLM Evaluation" (2026). Master's Projects. 1816.
DOI: https://doi.org/10.31979/etd.nfqd-nzvn
https://scholarworks.sjsu.edu/etd_projects/1816