APEXA: Execution-Integrity Enforcement for Multi-Agent LLM Automation of Synchrotron Data Reduction
LLM agents can make expert scientific workflows accessible through natural language, but a plausible response does not establish that the requested computation ran. This matters at synchrotron facilities, where calibration and integration are multi-step operations over terabyte-scale datasets. We present APEXA, a deployed framework coordinating 61 tools for synchrotron data reduction. A deterministic execution-integrity guard prevents it from returning results for tool calls that did not execute, and the same tool layer enforces motor safety: across 50 adversarial scenarios and four models, it produced 0/200 violations against a simulated EPICS controller, versus 15/200 for a prompt-only baseline. We also introduce APEXA-Bench, 58 facility tasks organized by scientific correctness and physical consequence. On real APS beamline data, one prompt recovered detector geometry and integrated a full attenuation and exposure sweep. Trustworthy scientific agents need deterministic validation of execution, not model-side assurances. We release the framework, benchmark, and traces.