arXiv · 2609.32617
The Alignment Paradox: How Post-Training Amplifies Confident Hallucinations in Language Models
Abstract
Large language models (LLMs) can produce factually incorrect answers with high confidence, undermining their reliability and limiting the effectiveness of uncertainty-based error detection. While prior research attributes confident hallucinations to factors such as missing knowledge in training data, reasoning errors, or stochastic decoding, we uncover that post-training alignment itself is a primary driver of these errors, a phenomenon we call the \textbf{Alignment Paradox}. Across five model families evaluated on factual benchmarks, unaligned base models produce few high-confidence errors on long-tail factual queries, whereas instruction-tuned models multiply high-confidence errors ($p \ge 0.95$) by more than an order of magnitude (10$\times$ to 35$\times$). Layer-wise probing with the Logit Lens reveals that this overconfidence emerges in late layers, where wrong-answer margins expand past 4.0 points after remaining near zero across early and intermediate layers. These findings motivate limiting margin growth during post-training. We implement this principle through an entropy-dependent margin bound in direct preference optimization (DPO). In multi-epoch experiments with Mistral-7B, the bounded objective reduces high-confidence errors by up to 35.3\% relative to standard DPO while maintaining performance on evaluated general reasoning benchmarks. These results show that bounded margins mitigate confident hallucinations during post-training.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Qingjia Huang, Yakai Li, Jianguo Wu, Qihang Zhou, Aimin Yu, Xiaoqi Jia, Luping Ma, Weijuan Zhang. 2026-09-26. The Alignment Paradox: How Post-Training Amplifies Confident Hallucinations in Language Models. https://arxiv.org/abs/2609.32617
Cite the original work for its findings. Save a collection to share your selection of sources.