Search arXivSearch

arXiv subjects

Li

Publications and source records attributed to Li.

2 recordsLinked to original sources

READY or Not: Reliable Enterprise Agent Deployment

An AI agent can perform well on benchmarks and still be unsuitable for deployment. Existing AI-agent benchmarks measure whether an agent can complete realistic professional work, whereas enterprise deployment asks a different question: whether an agent can meet a required reliability level, under acceptable human oversight, and at tolerable cost. We introduce Reliable Enterprise Agent Deployment (READY), a framework for qualifying AI agents for deployment on enterprise workflows. READY preserves each workflow's own definition of successful execution while applying a common qualification procedure. Given an agent, a workflow, and a class of candidate oversight policies, READY measures the reliability and operating cost of the human-AI system, selects the minimum-cost policy that satisfies a specified reliability target, and statistically qualifies it on held-out cases. The resulting deployment profile characterizes the supported operating point: reliability, human-oversight burden, and cost. READY is implemented as an open testbed that decouples workflow specification, execution, evaluation, and qualification, and runs on existing agent-evaluation infrastructure. In an end-to-end clinical-audit case study spanning 16 agent systems and 750 cases, READY reveals differences hidden by autonomous performance: two systems separated by only 0.3 points in autonomous accuracy (72.8% vs. 72.5%) require 39.2% versus 29.6% human review, respectively, to qualify at the same 76% reliability target under the evaluated oversight policy. READY thus shifts enterprise agent evaluation from how well can the agent perform the work? to under what conditions, and at what cost, can it be reliably deployed? By making those conditions explicit and statistically testable, READY provides a basis for comparing agent systems, setting oversight requirements, and making evidence-based deployment decisions.

cs.AI

A Bayesian Adaptive Spectral Surrogate Model for Efficient Probabilistic Optimal Power Flow Evaluation

This paper presents an adaptive stochastic spectral embedding (ASSE) method to solve the probabilistic AC optimal power flow (AC-OPF), a critical aspect of power system operation. The proposed method can efficiently and accurately estimate the probabilistic characteristics (e.g., mean, variance, median, and quantile-based metrics) of AC-OPF solutions while minimizing power losses. Based on estimated AC-OPF decisions (i.e., generator outputs), the confidence interval (CI)-based production cost index can be determined. Specially, an adaptive domain partition strategy is adopted to guide refinement domain selection and partition. The Bayesian compressive sensing-based coefficient calculation algorithm is integrated to enhance its performance. Numerical studies on modified IEEE 9-bus and IEEE 118-bus systems demonstrate that the proposed ASSE method offers accurate and fast evaluations compared to Monte Carlo simulations. Comparisons with a sparse polynomial chaos expansion, Gaussian process regression, and deep neural networks, further illustrate its efficacy in accurately assessing the responses with strongly localized behavior and non-symmetric distributions, providing practical decision-making bounds for generator outputs and operating costs under uncertainty.

eess.SY