PrefixGuard: Online Failure Warning and Trace-Grounded Diagnosis for LLM Agents
Large language model (LLM) agents execute multi-step tasks, but terminal verdicts arrive too late for intervention. Agent traces mix messages, tool calls, and feedback, making it difficult to learn failure signals from terminal outcomes. Even if a model learns to predict failure from execution prefixes, its predictions alone cannot show which actions, feedback, or unmet conditions need inspection. Existing work studies online failure prediction and diagnosis from completed trajectories, yet few methods connect timely warning to a compact symbolic model of execution behavior. We introduce \textbf{PrefixGuard}, a neuro-symbolic method that uses learned events to connect online failure warning with finite-state execution diagnosis. A gated recurrent unit (GRU) predicts horizon-specific failure risk from these events. A deterministic finite automaton (DFA) organizes their histories into compact risk-labeled paths for replay and rule-monitor composition. Across all 24 benchmark--horizon settings, PrefixGuard exceeds the strongest evaluated baseline in mean test area under the precision--recall curve (AUPRC). Average gains range from 12.0 to 21.9 percentage points. At a nominal 10\% FAR budget, averaged test recall spans 56.1\%--93.0\% across benchmarks. On one frozen $τ^2$-Bench DFA, model checking five specifications identifies violations in 21 of 265 benchmark-successful runs.