arXiv · 2610.05333
Diffusion Transformers are Provably Optimal In-context Generators
Abstract
Generative foundation models are attracting interest for their ability to produce desired outputs from demonstrations given at inference time, without updating parameters. However, since a few demonstrations cannot uniquely identify the intended task, the challenge is how to learn and sample from an output distribution that reflects this task uncertainty. In this work, we theoretically analyze how a Diffusion Transformer (DiT), pretrained across diverse tasks, learns and generates predictive distributions for a new query from demonstrations. We first show that the natural target to generate from finite demonstrations is not an output derived from estimating a single task, but rather a predictive distribution that captures the task uncertainty remaining after observing the demonstrations. We then prove that a DiT can learn this predictive distribution through score estimation, using attention to aggregate information from demonstrations and diffusion to generate samples. Owing to this property, with sufficient pretraining resources and diffusion sampling steps, the resulting DiT achieves the minimax optimal rate over a Hölder class of test-time tasks. These results imply that DiT acts as a statistically grounded in-context generator capable of generating distributions adapted to new tasks while retaining the uncertainty inherent in finite demonstrations.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Guoji Fu, Tomoya Wakayama, Ryotaro Kawata, Atsushi Nitanda, Wee Sun Lee, Taiji Suzuki. 2026-10-04. Diffusion Transformers are Provably Optimal In-context Generators. https://arxiv.org/abs/2610.05333
Cite the original work for its findings. Save a collection to share your selection of sources.