Causal Behavioral Evaluation of AI Agents at Scale via Automated Behavioral Science
As AI agents are increasingly deployed in complex and new environments, knowing the conditions that influence their behavior becomes an indispensable step for their reliable and safe deployment. Yet causal behavioral evaluation of AI agents remains manual and labor-intensive. We introduce Abs2Sim and AEROBAT, a system of methods that support causal behavioral evaluation of AI agents via automated behavioral science. Given a user-specified target behavior, the methods automatically execute a full pipeline of behavioral science research---generating hypotheses about the behavior, designing and executing controlled experiments, making behavioral assessments, analyzing the results, and writing reports. For 12 target behaviors, the methods generated and tested 73 hypotheses: designing 1,160 controlled experiments and executing 22,954 simulation rounds in total. Moderate-to-strong statistical evidence emerged for 30 hypotheses, revealing potential modulators of AI behavior. In sum, our results demonstrate that automated behavioral science can extend the reach of behavioral evaluation of AI agents.