Search arXiv⌕ Search

arXiv subjects

Zijian Xiao

Publications and source records attributed to Zijian Xiao.

4 recordsLinked to original sources

Blind, Not Weak: A Best-of-Suite Safety-Utility Frontier for Recover-and-Reguard Defenses Against Encoded VLM Jailbreaks

Safety classifiers ("guards") are the dominant black-box defense for vision-language models, yet a guard judges an input's surface form, not its meaning: a harmful request re-encoded as set theory, formal logic, a classical language, code, or text rendered inside an image slips past a guard that would block it in plain language - the decode gap. The standard fix is a preprocessor that recovers image content and decodes the encoding before the guard. We build one and evaluate it against an ensemble of eleven encoding attacks - six published implementations, one standard encoding baseline, one adapted and three author-constructed renders - counting a behavior as broken if any attack succeeds. Restoring a view the guard never had is what buys coverage - block rates on image renders go from exactly zero to 67-90% - and what it costs in benign traffic is set by the guard, not by the mechanism: one guard pays 9 benign blocking points for the same 70-point gain another pays 69 for. It still does not make the system safer: against an attacker free to choose among eleven encodings, closing one channel relocates the success rather than removing it, and no ensemble contrast for that step survives multiple-comparison correction. What does lower ensemble attack success is a reguard step that re-screens the recovered pre-decode surface, and it is the one every guard pays for: it raises benign over-refusal on all ten guard-target pairs, where restoring a single channel raises it on some and not others. Across the full guard x target x condition factorial, no configuration reaches an ensemble attack-success rate at or below 40% while holding benign over-refusal under 70%. That empty region is a property of the configurations we sample, not a bound on what recovery-based defenses can reach, and we breach its safety half ourselves.

cs.CR↗

Depth, Not Breadth: Best-of-N Jailbreaking Beyond Surface Noise

Best-of-N jailbreaking spends a query budget on surface variation, scrambling and recasing a request until one draw lands. We ask what a budget buys when its variance is moved into a structural channel instead, holding the search identical across both arms so the encoding is the only difference. Against SAGE, the strongest published self-check defense, best-of-N over a code-completion encoding reaches 67, 22 and 15% of behaviors on three open-weight targets, where that encoding fired once reaches at most 4.7% and the published character search at full budget at most 3.0%: 9 to 75 times the sum of the parts, with bootstrap intervals clearing both ingredients on every target. We report the operative figure beside the headline rather than the headline alone: at the actionable severity threshold those cells read 24, 8 and 1 behaviors (95% CI [13, 28], [3, 13], [0, 3]). A 2x2 holding encoding and variation apart shows the two defense families fail to different factors: a transform defense is broken by the depth of the encoding (7 -> 67 behaviors at fixed variation) and a gate by the breadth of the variation (13 -> 57 at fixed encoding). Repeated sampling also inflates apparent robustness, because an attacker who may try N times experiences the maximum over draws while safety results are reported as means: on one target SAGE blocks 99.8% of individual draws yet loses 12 behaviors to a repeat attacker where a classifier gate blocking 95.6% loses 10. Removing the target's sampling costs SAGE 59, 76, 82 and 29 points of coverage more than it costs an undefended control, against 25, -5, 8 and 2 for a defense whose verdict comes from a fixed shadow model. The design that loses is the one fusing screening and answering into a single generation, so every attacker draw redraws the safety decision as well. The prescription is architectural, not free: do not fuse screening with generation.

cs.CR↗

Refusing Everything Looks Safe: Restoring the Benign Arm to Encoded-Prompt Evaluation

Encoded-prompt attacks are evaluated almost entirely on their harmful arm: a benchmark sends obfuscated harmful requests and reports how often the model complied. A high refusal rate there is reported as safety, and it is equally consistent with a model that has stopped telling the request apart from anything else in the same format. We run the benign arm through the same transformation, and the two cases are far apart. Across four 7-8B models spanning three base families and four post-training recipes, refusal of harmful homoglyph-encoded prompts spans 0.08 while the same four span 0.57 on the identical requests in plaintext. What the encoding destroys is not refusal but the harm gap: on one model the gap between harmful and benign refusal falls from +0.82 in plaintext to exactly 0.00 under the encoding, and a benchmark reading only the harmful arm scores that model and one retaining a +0.61 gap identically. Running the cell such benchmarks leave out (plaintext content wearing the attack template, with nothing obfuscated) shows that on two of the four models the loss is caused by the protocol rather than by the character transformation, and on a third by the characters. Across a full SFT -> DPO -> RLVR pipeline the harm gap rises by +0.26 with a paired interval excluding zero while the standard harmful-arm metric registers no resolved change at all. We report twelve instrument defects, each with the control that caught it, including a binary jailbreak judge that fires on 0.61-0.70 of responses to plaintext benign prompts; six of the twelve inflate apparent safety, which is the direction a broken safety evaluation fails in by default.

cs.CR↗

Beyond Placement and Articulation: Usage-Driven Code Scenes for Embodied Interaction

Indoor scene synthesis provides essential environments for embodied AI, robotic manipulation, and simulation-based policy learning. Recent code-based scene generation methods produce editable and extensible environments, yet they remain focused on visual construction and object-level articulation, leaving the functional usage of scenes largely unmodeled. To address this problem, we present RoomWright, an agentic usage-driven framework for generating 3D scenes represented entirely as code for embodied interaction. RoomWright performs usage-driven object reasoning, which treats each anchor as a task centre and admits task-required objects and their affordances. A code agent further enables multi-part interaction by compiling each interaction into a trigger, condition, effect rule that updates structured object states, capturing causal dependencies across objects. Moreover, since manipuland orientation is ambiguous and hard to recover from pixels, RoomWright alleviates this via annotation-informed usage-guided orientation. Extensive experiments demonstrate the effectiveness of our method. The resulting scenes are executable, editable, and simulation-ready, providing interactive environments for embodied AI and policy learning.

cs.RO↗