arXiv · 2609.37670
MeanFlowAdvantage: Stable Reward Fine-Tuning for Few-Step Average-Velocity Generators
Abstract
MeanFlow enables efficient few-step generation by predicting interval-average velocities, but this representation creates a mismatch for reward fine-tuning: existing advantage-based objectives are typically defined on instantaneous velocities or equivalent $x_0$-space predictions, whereas inference directly uses the learned average-velocity map. We introduce MeanFlowAdvantage, a signed advantage-weighted least-squares objective for average-velocity generators. Our key construction uses a shared, detached MeanFlow derivative correction to express the reward objective in prediction space while making rollout and reference regularization exact penalties on the average-velocity network deployed at inference. The resulting formulation preserves MeanFlow's native few-step sampler and provides a direct mechanism for transferring reward improvements to the deployed flow map. On SD3.5-Medium, MeanFlowAdvantage improves all eight reported metrics over the matched four-step MeanFlowNFT baseline and, with only four NFEs, matches or exceeds the 40-step DiffusionNFT baseline on six of eight metrics. The same objective also transfers to DNA promoter design, where it supports both teacher-free on-policy RL for a generator defined on a manifold and teacher-guided reward-graded distillation, with the latter yielding the lowest one-step Sei profile MSE among the compared configurations.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Haocheng Tang, Tianchi Xie, Xingqiao Lin. 2026-09-29. MeanFlowAdvantage: Stable Reward Fine-Tuning for Few-Step Average-Velocity Generators. https://arxiv.org/abs/2609.37670
Cite the original work for its findings. Save a collection to share your selection of sources.