arXiv · 2609.29697
Understanding and Exploiting Initialization Anchoring Weakness in Feedback-Based Agent Planning
Abstract
Feedback-based planning improves agent reliability by incorporating tool observations and corrective feedback. However, its protection may not be distributed uniformly across planning stages. We conduct a round-wise analysis of four representative feedback mechanisms and uncover an initialization anchoring weakness: the first feedback round corrects 46\% of adversarial directions, whereas the rates fall to 13\% and 7\% among directions surviving into the next two rounds. Our analysis attributes this weakness to three interacting factors: a contextually plausible shift in the initial plan, insufficient counterevidence, and the persistence of accepted directions in the accumulated trajectory. Based on these findings, we propose \textsc{InitAnchor}, a black-box framework for exploiting this weakness through attacker-controlled external materials. It operationalizes the three factors as directional-shift, contextual-plausibility, and counterevidence-resilience signals under either limited target access or no target access. Across 112 tasks from 16 domains, six agent architectures, and five backbone LLMs, \textsc{InitAnchor} achieves average ASRs of 76.1\% and 72.0\% under the two settings while reducing first-round mitigation rates to 21.0\% and 25.0\%, respectively. It also remains effective against six defenses and across six real-world agent systems. These findings show that feedback-based agents can retain early biases even when later correction is available.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Chuanchao Zang, Jianing Wang, Wenyu Chen, Xiangtao Meng, Li Wang, Xinyu Gao, Peng Zhan, Zheng Li, Shanqing Guo. 2026-08-31. Understanding and Exploiting Initialization Anchoring Weakness in Feedback-Based Agent Planning. https://arxiv.org/abs/2609.29697
Cite the original work for its findings. Save a collection to share your selection of sources.