arXiv · 2609.38962
Alleviating Hallucination in Reasoning Tasks with Training-Free Uncertainty-Guided Steering
Abstract
Recent work on hallucination detection in large language models has shown that, for a fixed pre-trained model and reasoning task, it is possible to estimate the model's confidence in the correctness of its outputs. Such uncertainty estimates have primarily been used to improve truthfulness by detecting or filtering confabulations. In this work, we ask whether these signals can instead be used more proactively to directly improve the accuracy of model-generated answers. We propose USteer, a simple, training-free steering mechanism that adjusts a model's layer-wise activations during inference using the gradient of a confidence measure with respect to the activations. This procedure nudges generation toward outputs with lower uncertainty at inference time, without modifying model parameters or requiring additional supervision. We show that this approach consistently reduces hallucination across a range of tasks, demonstrating that confidence signals can be leveraged not only for detection, but also for effective inference-time control of model behavior.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Litian Liu, Qiqi Hou, Yubing Jian, Reza Pourreza, Mohammad Ghavamzadeh, Roland Memisevic, Yao Qin, Hong Cai. 2026-09-30. Alleviating Hallucination in Reasoning Tasks with Training-Free Uncertainty-Guided Steering. https://arxiv.org/abs/2609.38962
Cite the original work for its findings. Save a collection to share your selection of sources.