Unmodeled states and uncertain action outcomes in agentic scanning tunneling microscopy
In physical experiments, interventions can alter hidden experimental states in ways that cannot be predicted in advance. Autonomous scientific agents must therefore interpret the consequences of their interventions while operating with incomplete knowledge of the experimental state. Here we investigate this problem using scanning tunneling microscopy (STM) tip conditioning, traditionally a human-expert-intensive task governed by inaccessible tip apex conditions and uncertain action outcomes. We introduce quailbot, an agent harness that places the LLM inside the instrument feedback loop by linking physical interventions with experimental readbacks. After a brief apprenticeship with a human expert, frontier LLM agents autonomously completed end-to-end tip conditioning on an STM and passed an independent verification. The contingencies that arose during the experiments further exposed the limits of LLM agentic autonomy when relevant experimental states or dynamics were hidden or unmodeled. Our results show that future autonomous physical experimentation requires agents to infer hidden and unmodeled experimental states, track the outcomes of their actions, and operate within the observability and action limits of real instruments.