AeroCopilotBench: Safety-Gated Evaluation of LLM Agents on Aircraft Emergency Procedures in an Executable Cockpit
Aviation knowledge question answering cannot directly assess the operational effectiveness and safety compliance of large language models throughout aircraft emergency procedures. We introduce AeroCopilotBench and its executable cockpit environment, ACOE, which define state-transition rules, task goals, and trajectory-level safety constraints based on aircraft-specific Pilot's Operating Handbooks (POHs). The benchmark comprises 12 scenario templates and 73 tasks across two aircraft types, evaluating task completion, safety compliance, execution discipline, and repeatability. Across repeated evaluations of 12 models, the highest safety-gated success rate is 72.6\%. Most failed episodes achieve all critical goals but do not satisfy all remaining terminal goals. Across repeated runs, models still omit steps that they execute in other runs of the same task. Analysis of failed trajectories further reveals that some reasons for actions that conflict with the aircraft's POH recur across models. These results expose shortcomings in complete procedure execution, consistency across runs, and aircraft-specific emergency response.