Deployment of AI-Assisted Interventions: Capacity Constraints and Noisy Compliance
AI tools increasingly drive targeted interventions in various service settings, including healthcare, education, and public services. Algorithms score individuals, trigger outreach to those above a threshold (e.g., high-risk or high-value), and encourage them to request service; then providers deliver service to those who request. Much of the work in this area has focused on improving predictive accuracy, implicitly assuming that better predictions lead to better outcomes. We show that predictions are only one component of a larger service system: when service capacity is limited and behavioral responses to outreach are probabilistic, system efficacy depends on operational forces that predictive accuracy does not capture. In such settings, the optimal score threshold must balance two effects: ensuring all capacity is filled (utilization) and, when capacity is constrained, preventing low-value requests from crowding out high-value ones (cannibalization). We characterize the optimal threshold and prove that thresholds based solely on predictive accuracy are generally suboptimal. Further, algorithm selection metrics such as AUC can be misaligned with operational performance: they weight all thresholds equally, while optimal deployment uses a subset of thresholds that depends on both capacity and compliance behavior. We introduce a new metric, Operational AUC (OpAUC), and show that it identifies the efficacy-optimal algorithm. Finally, we conduct a case study on sepsis early warning data that illustrates the magnitude of the improvement available from better threshold selection and shows that a predictor with lower AUC can achieve higher system efficacy under optimal deployment.