ESCROW: Guarded and Dual-Objective Continual Maintenance for Agents in Policy-Governed Enterprise Workflows
LLM agents increasingly run policy-bound enterprise workflows, where they must apply rules consistently and stay auditable. Deploying such an agent is the start of its long-term maintenance cycle: it must adapt to a stream of operational signals, yet reliably turning these sparse, unlabeled signals into reusable skill revisions is hard, and a careless update can trade one task category's accuracy for the overall gain, revive a resolved failure, or land at an undeployable cost. We present ESCROW, a post-deployment maintenance framework that updates an agent's external, reviewable skills under a Strict Update Boundary: the LLM proposes candidate revisions, but only an empirically evaluated version is deployed. It combines distributed diagnosis with consensus, a per-category non-regression guard, cross-cycle anti-regression, and accuracy--cost Pareto search, emitting a versioned, auditable diff per change. In real production on our internal financial document-auditing system, it attains the strongest evaluated accuracy--cost trade-off among baselines, with a transfer probe on public $τ$-bench.