arXiv · 2609.33702
Understanding Confabulation and Rethinking Reconstruction in Activation Explanations
Abstract
Natural Language Autoencoders (NLAs) produce unsupervised text explanations of a model's activations: a verbalizer describes an activation and a reconstructor learns to recover it from this text. Under the established point-reconstruction NLA training recipe, explanations become more useful for predicting model behavior while also increasingly introducing unsupported details and exhibiting writing defects. To assess these changes separately, we introduce a standardized evaluation framework for unstructured NLA explanations, measuring information recoverable from explanations, contextual support for their claims, and writing quality. To address confabulation and writing defects, we move beyond predicting a single activation: explanations can distinguish distributions of possible activations even when their means and optimal point-reconstruction rewards are identical. We introduce Flow-NLA, which models the distribution of activations compatible with an explanation and trains the verbalizer using a diffusion likelihood bound. Across Qwen, Gemma, and Apertus, this richer signal retains the utility gains of point reconstruction while curbing the growth of confabulation and writing defects, opening up a direction for improving activation-derived training to encourage more informative, supported, and readable explanations. Code and evaluation prompts will be made publicly available upon acceptance.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Gert Lek, Zixuan Xia, Pin-Yu Chen, Lydia Y. Chen. 2026-09-27. Understanding Confabulation and Rethinking Reconstruction in Activation Explanations. https://arxiv.org/abs/2609.33702
Cite the original work for its findings. Save a collection to share your selection of sources.