arXiv · 2610.05572
Disentangling Task Difficulty from Run-Level Failure in Agent Failure Prediction
Abstract
Predicting whether an LLM agent will fail has emerged as a promising direction for supporting intervention during execution. Recent approaches report strong predictive performance, often with AUROC values between 0.85 and 0.94. However, predictors are typically trained by pooling runs from many tasks. We hypothesize that part of this performance comes from recognizing that some tasks are harder than others, rather than detecting whether a particular run is heading toward failure. This distinction matters because task-level difficulty supports decisions about where to allocate computation, while run-level prediction is needed to decide whether to intervene in an ongoing trajectory. We study benchmarks with repeated attempts of the same task by the same agent and separate cross-unit comparisons from comparisons between successful and failed runs of the same model-task unit. Across the evaluated corpora, more than 99.93% of the positive-negative pairs underlying pooled AUROC are cross-unit. Accordingly, predictors that never observe the current run can achieve high pooled performance, including a difficulty oracle with AUROC up to 0.945, while remaining at chance within task. Early run-level discrimination is consistently weak across trajectory predictors, released monitors, and hidden-state probes, although it improves later in execution and is stronger for weaker agents. Under fixed token budgets, task-level allocation outperforms abort-only strategies, while early stopping becomes beneficial only when within-task AUROC reaches about 0.84-0.93, far above the 0.50-0.55 range observed for early monitors. These results show that failure prediction should be evaluated not only by outcome accuracy, but by whether the captured signal supports the intended deployment decision.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Mohsen EsfandyariDoulabi, Lawrence Arkoh, Biruk Tadesse, Vaishvi Patel, Mehul Sharma, Marcelo d'Amorim, Wesley Assunção. 2026-10-04. Disentangling Task Difficulty from Run-Level Failure in Agent Failure Prediction. https://arxiv.org/abs/2610.05572
Cite the original work for its findings. Save a collection to share your selection of sources.