arXiv · 2604.23478
JudgeSense: A Benchmark for Prompt Sensitivity in LLM-as-a-Judge Systems
Abstract
Large language models are widely used to judge the output of other language models, yet whether a judge returns the same verdict when the same request is worded differently remains largely unexamined. We study that question across four evaluation tasks and twenty-five judges from six providers. To support the analysis we release JudgeSense, a benchmark of 880 items from human-labelled corpora, each issued under two instructions that differ in wording and not in what they ask, with the complete decision logs. Every score is reported against the judge's own agreement with itself on the identical prompt, so decoding noise is not charged to wording, and the release lets a reader ask the same of any judge not in our roster. Rewording costs agreement on all four tasks, and on two it clears the threshold we declare for a practically meaningful effect; the ordinal task is both the least stable and the one fewest judges are accurate on, and within a single family parameter count does not predict stability. A judge measured inside an agent harness yields a smaller estimate than the same judge reached through a direct API call, because its agreement with itself collapses faster than its agreement across wordings.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Rohith Reddy Bellibatlu, Edward Raff, Wenbin Zhang. 2026-09-18. JudgeSense: A Benchmark for Prompt Sensitivity in LLM-as-a-Judge Systems. https://arxiv.org/abs/2604.23478
Cite the original work for its findings. Save a collection to share your selection of sources.