Search arXiv⌕ Search

arXiv · 2610.07723

The Model Plants the Trigger: Answer-Side Backdoor Attacks in Multi-Turn Large Language Models

Abstract

Safety alignment in Large Language Models (LLMs) remains vulnerable to backdoor attacks. Existing LLM backdoors are almost all input-centric: activation depends on explicit trigger patterns in the user input, so modern guardrails are built to sanitize the input space. We challenge this assumption with a novel answer-side backdoor for multi-turn dialogue. Instead of inserting the trigger into the input, the adversary uses a benign first-turn prompt to naturally induce the model to generate a specific, seemingly innocuous word. Once merged into the dialogue history, this self-generated word becomes the trigger. When a later harmful query arrives, the model detects its own trigger and bypasses its safety refusal, while the user input stays perfectly clean. Across four LLMs, our attack reaches near-perfect Attack Success Rates, approaching 100\% at only a 5\% poisoning rate, while preserving general utility and clean-input safety, and it evades mainstream input-centric defenses. Representation-level analysis shows that the self-generated trigger consistently suppresses the model's refusal signal, exposing a critical blind spot in current LLM defenses.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Yibo Zhang, Tianrong Guan, Liang Lin, Puze Wang, Jin Wang, Qingsong Wen. 2026-10-06. The Model Plants the Trigger: Answer-Side Backdoor Attacks in Multi-Turn Large Language Models. https://arxiv.org/abs/2610.07723

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Adversarial Images Hijack Web Agents from Visual Grounding to Browser Execution

Modern web agents built on large vision-language models process webpages, select relevant UI elements, and translate model outputs into browser actions. Existing visual red-teaming approaches use adversarial visual content to manipulate this process. However, they primarily target model inference and do not explicitly account for structured input processing or action post-processing. Consequently, model-level success does not establish control over browser execution and cannot reliably characterize end-to-end agent robustness. To address this gap, we formulate red teaming for vision-grounded web agents as an end-to-end grounding-to-execution problem, and introduce WebMirage, a framework that crafts localized visual perturbations that cause agents to select attacker-controlled content and execute the corresponding browser action across varying webpage renderings. It uses a role-slot abstraction and webpage recomposition to capture competition among webpage elements, and dataflow analysis to align optimization with action post-processing. We evaluate WebMirage across four agent configurations and six VLM backbones on 2,250 tasks covering 13 public websites and a sandbox benchmark. WebMirage achieves an average attack success rate of 91.9%, compared with 17.4% for the strongest baseline, and remains effective against three agent-level defenses.

cs.CR↗

ICS-Sniper: A Targeted Blackhole Attack on Encrypted ICS Traffic

Modern industrial control systems (ICS) increasingly host their Supervisory Control and Data Acquisition (SCADA) services in the cloud to reduce the costs of large-scale automation. To protect site-SCADA communications, ICS operators commonly use VPN tunneling and standard security practices. We show that, despite these security measures, an on-path Internet adversary can disrupt ICS operations without infiltrating the ICS perimeter, breaking encryption, or knowledge of the control logic. We present ICS-Sniper, a targeted blackhole attack that analyzes the VPN traffic metadata (sizes, direction, timing of packets) to identify narrow time windows, called critical superperiods, during which the site-SCADA traffic would likely contain highly critical commands or data. Post-analysis, in a subsequent operational cycle, ICS-Sniper drops a small set of payload-carrying packets in the critical superperiods to disrupt the ICS's operations. We demonstrate three attacks on two realistic modern Secure Water Treatment (SWaT) plant testbeds that can potentially violate the operational safety of the ICS while evading state-of-the-art ICS attack detectors.

cs.CR↗

Large Language Models for Cryptocurrency Transaction Analysis: A Bitcoin Case Study

Cryptocurrencies are widely used, yet current methods for analyzing transactions often rely on opaque, black-box models. While these models may achieve high performance, their outputs are usually difficult to interpret and adapt, making it challenging to capture nuanced behavioral patterns. Large language models (LLMs) have the potential to address these gaps, but their capabilities in this area remain largely unexplored, particularly in cybercrime detection. In this paper, we test this hypothesis by applying LLMs to real-world cryptocurrency transaction graphs, with a focus on Bitcoin, one of the most studied and widely adopted blockchain networks. We introduce a three-tiered framework to assess LLM capabilities: foundational metrics, characteristic overview, and contextual interpretation. This includes a new, human-readable graph representation format, LLM4TG, and a connectivity-enhanced transaction graph sampling algorithm, CETraS. Together, they significantly reduce token requirements, transforming the analysis of multiple moderately large-scale transaction graphs with LLMs from nearly impossible to feasible under strict token limits. Experimental results demonstrate that LLMs have outstanding performance on foundational metrics and characteristic overview, where the accuracy of recognizing most basic information at the node level exceeds 98.50% and the proportion of obtaining meaningful characteristics reaches 95.00%. Regarding contextual interpretation, LLMs also demonstrate strong performance in classification tasks, even with very limited labeled data, where top-3 accuracy reaches 72.43% with explanations. While the explanations are not always fully accurate, they highlight the strong potential of LLMs in this domain. At the same time, several limitations persist, which we discuss along with directions for future research.

cs.CR↗