arXiv · 2609.34497
QAMM: Adjoint MeanFlow Matching for Few-Step Offline Reinforcement Learning
Abstract
Flow policies can model rich action distributions, but their iterative sampling limits decision speed. Adjoint matching uses the critic's action gradient to improve a flow policy without backpropagating through its sampling trajectory, yet its supervision is defined for instantaneous velocities. We propose QAMM, a method that turns the critic-derived adjoint signal into supervision for MeanFlow's average velocity. The resulting policy learns finite-interval transport directly and generates actions with few network evaluations. We derive the adjoint MeanFlow target, specify its gradient boundaries, and train it with an offline actor-critic. On ten HumanoidMaze tasks, QAMM produces effective two-call policies and achieves competitive performance against strong flow-policy baselines. These results show that adjoint-based Q optimization can be combined with average-velocity learning to obtain expressive offline policies with few-step action generation.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yuehu Gong, Shutong Ding, Mokai Pan, Yimiao Zhou, Jiashu Hou, Ye Shi, Yanwei Fu. 2026-09-28. QAMM: Adjoint MeanFlow Matching for Few-Step Offline Reinforcement Learning. https://arxiv.org/abs/2609.34497
Cite the original work for its findings. Save a collection to share your selection of sources.