arXiv · 2610.05563
G-CARB: Graph-Localized Conformal Agent Risk Budget for Compositional Harm
Abstract
Small language model (SLM) agents need safety controls that track consequences across tool calls with little monitoring overhead. A private read, for example, becomes a leak when a later action sends that data outside the system. We introduce CARB (Conformal Agent Risk Budget), which calibrates when to stop an agent using a ledger of harm incurred before stopping. Under exchangeable episodes, standard conformal risk control bounds this declared loss in expectation over calibration and a future episode. G-CARB selects scorer evidence along observable dependencies from private sources to outgoing actions. The ledger still covers the entire executed history, and computing the gate score requires no additional language-model inference. On AgentDojo replay with two 14B backbones, G-CARB roughly halves scorer-input records at intermediate risk budgets while improving autonomous task completion relative to full-prefix scoring; random context of the same size achieves similar gains. Controlled examples show how retaining the relevant dependency can further avoid stopping benign work.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Zijun Yu, Yu Gu, Vahid Partovi Nia, Masoud Asgharian. 2026-10-04. G-CARB: Graph-Localized Conformal Agent Risk Budget for Compositional Harm. https://arxiv.org/abs/2610.05563
Cite the original work for its findings. Save a collection to share your selection of sources.