arXiv · 2609.33444
Elucidating the Design Space of Regression-based Diffusion Reinforcement Learning
Abstract
A nascent family of methods that forgoes the policy gradient and reweights a supervised regression instead has garnered momentum in reinforcement learning for diffusion and flow models. DiffusionNFT, FlowAWR, and RAM are representative regimes with contrasting motivations. It is yet opaque what, if anything, they share. We substantiate that each is the solution of one divergence-constrained reward-maximization problem, and they are differentiated only by the convex generator that defines the constraint. Under the unified modeling framework, we unravel the relaxations that prior art made during building the advantage-embedded regression target: approximating the KKT condition and posterior normalizer for the linear and exponential tilt shapes DiffusionNFT and FlowAWR respectively, while preserving the exact sparsemax projection onto the probability simplex for linear tilt leads to another superior model type in this work. Beyond the theoretical underpinnings, we further empirically investigate the design space and shed light on the training recipe for regression-style diffusion RL. Retaining the merits discovered during our exploration gives rise to DiffusionRFT, our paradigm that converges faster, trains more stably, and attains the top performance.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Toyota Li, David Zhao, Alan Zhao. 2026-09-27. Elucidating the Design Space of Regression-based Diffusion Reinforcement Learning. https://arxiv.org/abs/2609.33444
Cite the original work for its findings. Save a collection to share your selection of sources.