Search arXivSearch

arXiv subjects

L. Degeorge

Publications and source records attributed to L. Degeorge.

1 recordsLinked to original sources

How far can we go with ImageNet for Text-to-Image generation?

Recent text-to-image (T2I) generation models have achieved remarkable sucess by training on billion-scale datasets, following a `bigger is better' paradigm that prioritizes data quantity over availability (closed vs open source) and reproducibility (data decay vs established collections). We challenge this established paradigm by demonstrating that one can achieve capabilities of models trained on massive web-scraped collections, using only ImageNet enhanced with well-designed text and image augmentations. With this much simpler setup, we reach the performance of FLUX and achieve a +5 overall score over SD3 on GenEval and +12 on DPGBench over SDXL while using just 1/1000th the training images and 3x to 10x less parameters. This opens the way for more reproducible research as ImageNet is widely available and the proposed standardized training setup only requires 500 hours of H100 to train a text-to-image model.

cs.CV