arXiv · 2602.03422
RankSteer: Can Pointwise LLM Rankers Be Calibrated at the Representation Level?
Abstract
Large language models (LLMs) are strong zero-shot pointwise rankers, but lag behind pairwise and listwise methods. Beyond missing comparative signals, we identify a \textit{calibration gap}: ranking-relevant information encoded in hidden states is not fully captured by the scalar output head. We propose RankSteer, a post-hoc activation-steering framework that calibrates ranking via projection-based interventions along multiple directions at inference time: decision, evidence, and, optionally, role. This is achieved without updating model weights or introducing cross-document comparisons. We instantiate RankSteer on two structurally distinct pointwise variants and observe improvements over their respective baselines on most TREC DL and BEIR datasets across three backbones. This suggests that the calibration gap is a general property of pointwise rankers. Our additional geometric analysis shows that steering improves ranking by concentrating each query's document representations along an existing ranking geometry, offering new insight into how LLMs internally represent and calibrate relevance judgments.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yumeng Wang, Catherine Chen, Suzan Verberne. 2026-09-18. RankSteer: Can Pointwise LLM Rankers Be Calibrated at the Representation Level?. https://arxiv.org/abs/2602.03422
Cite the original work for its findings. Save a collection to share your selection of sources.