arXiv · 2609.36434
The Safety Operator: Modulating the Expression of Safety Instructions via Spectral Optimization
Abstract
Context tokens in a transformer-based language model can be absorbed into the model's weights as a multiplicative operator. We study this operator in the setting of safety instructions and show that influencing its dominant eigenvalue modulates how strongly the instruction shapes generation. We derive a Contrastive Safety Loss with a suppression weight that controls the tradeoff between emphasizing the safety instruction on harmful queries while suppressing it on harmless queries. Varying the suppression weight maps a relationship between the attack success and the over-refusal rates, supporting the hypothesis that the operator's eigenvalue acts as a continuous dial for the instruction's influence. Moreover, this relationship holds relatively independently of how the Safety Loss is parameterized, yielding Pareto-improved safety instructions for appropriate values of suppression weight.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Benoit Dherin, Michael Munn, Xavier Gonzalvo, Adrian Goldwaser, Blaz Bratanic, Ananth Balashankar, Andrey Vlasov, Pinzhi Huang, Cecile Loge, Nicole Mitchell, Andre Fernandes, Trilok Acharya, Wendy Kan, Ziyue Wang, Hanna Mazzawi, Felipe Tiengo Ferreira, Mor Geva. 2026-09-29. The Safety Operator: Modulating the Expression of Safety Instructions via Spectral Optimization. https://arxiv.org/abs/2609.36434
Cite the original work for its findings. Save a collection to share your selection of sources.