arXiv · 2503.12018
CompArt: Operationalizing Aesthetic Alignment in Text-to-Image Generation via Principles of Art
Abstract
Text-to-Image (T2I) diffusion models have made rapid progress on semantic alignment (generating what is described in the prompt), yet users still lack reliable control over aesthetic composition (how visual elements are put together). Prior work often treats aesthetics as a single, preference-driven notion (e.g., "high quality", "detailed", "breathtaking"), which does not map cleanly to compositional intent. We propose Aesthetic Alignment: aligning generated images to explicit, user-specified compositional constraints. We operationalize these constraints using the Principles of Art (PoA)-e.g., Balance, Rhythm, and Emphasis-commonly used in art education to describe composition. To support this task, we introduce CompArt, a dataset of 80,032 WikiArt images augmented with captions and PoA analyses produced by a multimodal LLM under structured prompting. We further propose ArtDapter, a lightweight and disentangled adapter that enables steering a pretrained T2I model along 10 PoA dimensions while retaining the base model's semantic capability. Experiments on CompArt show improved adherence to PoA controls over strong baselines under a dual evaluation protocol.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Zhe Jin, Tat-Seng Chua. 2026-09-16. CompArt: Operationalizing Aesthetic Alignment in Text-to-Image Generation via Principles of Art. https://doi.org/10.1145/3805622.3810794
Cite the original work for its findings. Save a collection to share your selection of sources.