arXiv · 2609.36576
Divide and Inject: Can Agents Reconstruct an Indirect Prompt Injection from Fragments?
Abstract
Agentic systems are now being widely used to orchestrate tools and reason over long contexts. However, the improving capabilities of the large language models powering these agents also create new attack surfaces for indirect prompt injection. In particular, an attacker may not need to place a complete malicious instruction in retrieved content if the agent can reconstruct the objective from incomplete fragments distributed across a long context. In this work, we introduce adaptive long-context prompt injection (AdaLCPI), which combines long-context fragmentation with adaptive search. AdaLCPI splits an attack objective into incomplete fragments, embeds them in external content retrieved through the agent's tools, and uses a reconstruction cue to prompt the agent to combine them. It then iteratively refines the fragments and cue with OpenEvolve using graded scoring and natural-language execution feedback from the target agent. Empirically, AdaLCPI achieves higher attack success than strong adaptive baselines, reaching 61.4\% macro-average ASR compared with 32.8\% for Trojan Hippo-style and 30.0\% for AgentVigil. Safety evaluations should therefore test whether agents remain robust when harmful objectives must be reconstructed from incomplete fragments.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Michael Lee, Zhipeng Wei, Yue Dong, N. Benjamin Erichson. 2026-09-29. Divide and Inject: Can Agents Reconstruct an Indirect Prompt Injection from Fragments?. https://arxiv.org/abs/2609.36576
Cite the original work for its findings. Save a collection to share your selection of sources.