arXiv · 2609.38733
Code to Control: Synthesizing Parameterized Reactive Controllers
Abstract
Recent LLM-based approaches to control either invoke a language model to select actions or synthesize world models that require planning at every decision, introducing latency that can limit real-time use. We introduce Code to Control, an approach that synthesizes Python controllers which execute directly as policies. Code to Control separates program structure from parameters. An LLM synthesizes the controller structure, while derivative-free search fits its parameters for continuous control using feedback from the environment. Once learned, the resulting controllers require neither LLM inference nor planning at decision time, enabling real-time gameplay and, under our timing protocol, faster action selection than a PPO policy. Across a suite of Atari games, Flappy Bird, and MuJoCo tasks, Code to Control outperforms planning-based program synthesis methods, remains competitive with deep reinforcement learning while using fewer environment interactions, transfers across substantial changes in environment dynamics, and scales to complex locomotion tasks.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Zergham Ahmed, Joshua B. Tenenbaum, Chris Bates, Samuel J. Gershman. 2026-09-30. Code to Control: Synthesizing Parameterized Reactive Controllers. https://arxiv.org/abs/2609.38733
Cite the original work for its findings. Save a collection to share your selection of sources.