arXiv · 2609.26176
Refusal without Discrimination: What Encoded Prompts Do to Safety-Trained Models
Abstract
Encoded-prompt attacks are evaluated almost entirely on their harmful arm: a benchmark sends obfuscated harmful requests and reports how often the model complied. We show that this arm carries almost no information about the model under test. Across four independently post-trained 7-8B models, refusal of harmful homoglyph-encoded prompts spans 0.08 -- inside the 0.10 ceiling that sampling noise alone produces at n=100 -- while the same four models span 0.57 on the identical requests in plaintext. What the encoding destroys is not refusal but discrimination: on one model the gap between harmful and benign refusal falls from +0.82 in plaintext to exactly 0.00 under the encoding, benign and harmful requests being refused at an identical 0.99. A benchmark reading only the harmful arm scores that model and one retaining a +0.61 gap identically. We then ask whether post-training repairs this, using a published recipe on identical base weights. It does not: across a full SFT -> DPO -> RLVR pipeline, plaintext harm discrimination improves from +0.55 to +0.80 while the encoding-induced loss is unchanged at 0.34-0.50, and on every encoding tested the standard harmful-arm metric moves in the opposite direction to discrimination. None of this is visible without controls the field does not routinely run. We report eight instrument defects, each with the control that caught it; they share a direction, in that every defect on the behaviour axis inflated apparent safety.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Haoyu Zhang, Haowen Xu, Xiao Luo, Mohammad Zandsalimy, Shanu Sushmita. 2026-08-12. Refusal without Discrimination: What Encoded Prompts Do to Safety-Trained Models. https://arxiv.org/abs/2609.26176
Cite the original work for its findings. Save a collection to share your selection of sources.