C2P-VAR: Continual and Compositional Personalization in Visual Autoregressive Models
Visual autoregressive (VAR) models have recently emerged as an efficient paradigm for text-to-image generation, yet their personalization capabilities remain largely limited to static, single-concept settings. In practice, users may continuously introduce new concepts and wish to compose multiple personalized concepts within a single image. Such scenarios pose two fundamental challenges: catastrophic forgetting during sequential personalization and feature interference during multi-concept composition. In this work, we study continual and compositional personalization in VAR models and propose C2P-VAR, a unified framework addressing both challenges. For continual personalization, we introduce C2PVAR-S, which identifies concept-relevant parameters from gradient magnitudes and dynamically updates their selection during training. To preserve previously learned concepts, C2PVAR-S applies regularization only to parameters shared by the current and historical concepts, thereby reducing unnecessary interference without introducing additional model components. For multi-concept personalization, we further propose C2PVAR-M, which employs parallel global and concept-specific branches with spatially localized feature fusion and logit aggregation to achieve controllable concept placement and reduce feature entanglement. Extensive experiments on continual and multi-concept personalization demonstrate that C2P-VAR consistently outperforms existing baselines in subject fidelity, while maintaining competitive text alignment and introducing negligible storage and inference overhead. Our results establish a unified framework for scalable and controllable personalization of visual autoregressive models.