arXiv · 2609.27598
DriftAudio: Marginal Drifting for Distributional Post-Training of One-Step Text-to-Audio Generators
Abstract
Recent one-step text-to-audio (TTA) models substantially reduce inference cost, yet their generated distributions can still be improved through post-training. We propose DriftAudio, a distributional post-training method that adapts Drifting to pretrained one-step TTA generators. Applying Drifting condition-wise is challenging under free-form text conditioning, where only one or a few real samples are typically available for a particular condition. DriftAudio instead performs Drifting on the marginal audio distribution while retaining text-conditioned generation. The drifting field is estimated in a frozen audio feature space using resampled real samples, a rolling bank of generated samples, and the current batch of generated samples. The resulting Drifting field provides a detached training target for updating only the generator, keeping the original one-step inference procedure. On AudioCaps, starting from MeanAudio, DriftAudio reduces FAD and FD by 33.9% and 17.6%, respectively, while also improving KL and CLAP. Starting from FdAudio, it further reduces FAD, FD, and KL, with trade-offs in IS and CLAP. These results demonstrate the effectiveness of marginal distributional post-training for one-step TTA generation.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Xingyu Chen, Fei Ma, Sipei Zhao. 2026-09-23. DriftAudio: Marginal Drifting for Distributional Post-Training of One-Step Text-to-Audio Generators. https://arxiv.org/abs/2609.27598
Cite the original work for its findings. Save a collection to share your selection of sources.