arXiv · 2512.08463
Privileged observations enable rapid and reliable policy discovery directly in the physical world
Abstract
We study how privileged information about a physical system affects the discovery of high-performing policies when training a reinforcement learning agent directly in the physical world. We let the agent control a cylinder in a tabletop water channel to maximize or minimize drag. The flow is chaotic and difficult to model or simulate accurately and good strategies are not obvious beforehand. Decades-old experimental studies provide recipes for simple, high-performance, periodic open-loop policies that increase drag by 26.6% +/- 0.7% and reduce it by 29.7% +/- 1.3%. With dense observations of the cylinder wake, the agent learns within tens of minutes policies that increase drag by 25.5% +/- 0.9% and reduce it by 32.4% +/- 1.6%. We record action trajectories during online policy execution and replay them open loop as fixed sequences; these replays increase drag by 23.2% +/- 2.2% and reduce it by 32.1% +/- 3.0%. However, when we withhold flow observations during training, the agent still learns to decrease drag by 31.4% +/- 2.1%, but no run learns to increase it (1.8% +/- 4.3% drag increase). Our physical experiments demonstrate an extreme case: privileged observations can be decisive for policy discovery even when the resulting actions can be replayed in open loop.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Antonio Terpin, Raffaello D'Andrea. 2026-09-11. Privileged observations enable rapid and reliable policy discovery directly in the physical world. https://arxiv.org/abs/2512.08463
Cite the original work for its findings. Save a collection to share your selection of sources.