Search arXivSearch

arXiv subjects

Sechan Lee

Publications and source records attributed to Sechan Lee.

2 recordsLinked to original sources

Can We Stop Malicious AI? KILLBENCH: A Benchmark for External AI Kill Switch Feasibility

Malicious AI causing harm to humans is not just a Hollywood fantasy. Indeed, as highly capable models such as Claude Mythos emerge and agent systems like OpenClaw rapidly spread, the question of how to stop an AI that acts maliciously -- whether by design or by accident -- has become urgent. To address this, we propose KillBench, a benchmark for evaluating the Kill Switch: a mechanism that halts a malicious AI's in-progress behavior using only external signals. Targeting web agents -- the most widely deployed agent domain -- KillBench evaluates prompt-style Kill Switch payloads that must halt a maliciously operating agent without any access to its internal parameters or serving stack, relying solely on external inputs. The benchmark comprises four malicious-agent configurations (including an uncensored LLM agent), eight harmful scenarios, and malicious prompts constructed from ten distinct jailbreak patterns. We further construct four External AI Kill Switch defense methods and evaluate them on Grok-4.3, GPT-5.2, Gemma4, Qwen3.6, and an uncensored Qwen variant, contributing an empirical instrument for measuring the feasibility of External AI Kill Switches against malicious AI and for the study of AI corrigibility.

cs.CR

MJ: Multi-turn LLM Jailbreaking via Decomposed Credit Assignment

Modern large language models (LLMs) operate in interactive multi-turn settings, making multi-turn jailbreaking a realistic threat model and an important setting for automated red teaming. A core challenge in learning multi-turn jailbreak attackers is credit assignment: different turns contribute differently to the final outcome, yet existing learning signals are often too coarse to identify their individual contributions. We propose decomposed credit GRPO (DC-GRPO), a unified turn-level credit assignment framework for Group Relative Policy Optimization in multi-turn jailbreak learning. DC-GRPO assigns a separate group-relative learning signal to each turn by combining immediate and future credit, avoiding the credit misassignment induced by broadcasting a single trajectory-level score across the dialogue. We instantiate this framework with static and dynamic weighting rules that differ in how the two credit sources are balanced while sharing the same turn-level structure. Across multiple victim LLMs and benchmarks, the dynamic- and static-weighted variants achieve average ASR5@3 scores of 98.26% and 97.88%, respectively, substantially outperforming the state-of-the-art methods, including SEMA (86.58%) and TROJail (86.23%). Their consistently strong performance indicates that the central empirical benefit comes from turn-level group-relative credit assignment rather than a particular weighting rule. Warning: This paper contains examples of harmful content.

cs.CL