arXiv · 2610.02203
Embedding Prediction Helps Image Generation
Abstract
In diffusion transformers, a class label or a text prompt is embedded once, and the same condition is reused at every denoising step. We ask whether predicted embeddings can serve as this condition instead. Next-Embedding Predictive Autoregression (NEPA) trains a Transformer to predict the next continuous embedding in a sequence. In generation, the clean image follows the noisy image, so its embeddings are the next embeddings after the condition and the noisy image. We train a NEPA model to predict them all at once with Multi-Embedding Prediction, and in Embedding Conditioned Generation, a DiT generator is conditioned on these predictions, recomputed at every denoising step, so the conditioning signal adapts to the current noisy state. Experiments on class-conditional ImageNet $256\times256$ study the condition of the generator, the design of Multi-Embedding Prediction, and the scaling of both models. The NEPA model adds a second network to every sampling step; with it, and combined with REPA, our final model, NEPA-DiT-XL, reaches an FID of 1.32 using about a third of the training compute of REPA.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Sihan Xu, Ji Xie, Zilin Wang, Hui Shen, Stella X. Yu. 2026-10-01. Embedding Prediction Helps Image Generation. https://arxiv.org/abs/2610.02203
Cite the original work for its findings. Save a collection to share your selection of sources.