arXiv · 2602.00279
Semantic Calibration Prevails Where Token Confidence Fails: Benchmarking Long-Form Scientific QA
Abstract
Reliable uncertainty quantification (UQ) is essential for safe deployment of large language models (LLMs) in scientific question answering, where long-form outputs exceed practical human verification at scale. We introduce the first large-scale benchmark for UQ calibration in long-form, reasoning-demanding scientific QA, evaluating four UQ methods on 685,000 responses across up to 20 LLMs and seven datasets, supported by an extensible open-source framework whose shared-generation design enables reproducible cross-method comparisons. Instruction tuning is shown to associate with systematic token probability polarization, collapsing confidence distributions and undermining the reliability of token-level uncertainty signals. Reasoning model families diverge: some reproduce this polarization while others actively mitigate it, a pattern that clusters by provider and suggests training pipeline design as a key differentiating factor. Verbalized and token-aggregation sequence-level methods fail systematically. Only semantic consistency, as measured by consistency of the final answer, yields well-calibrated outputs, providing the first large-scale evidence that semantic calibration persists in multi-step, dependency-rich reasoning settings.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Philip Müller, Nicholas Popovič, Michael Färber, Peter Steinbach. 2026-09-18. Semantic Calibration Prevails Where Token Confidence Fails: Benchmarking Long-Form Scientific QA. https://arxiv.org/abs/2602.00279
Cite the original work for its findings. Save a collection to share your selection of sources.