Autonomous Fashion Outfit Composition via Unified Aesthetic Foresight Model
Fashion Outfit Composition (FOC) requires sequentially assembling fashion items into a stylistically cohesive ensemble. Existing works struggle to model this step-by-step process effectively, primarily because they fail to jointly optimize the two critical capabilities required for FOC: intermediate outfit value evaluation and complementary item prediction. This structural disconnect, compounded by the severe sparsity of step-wise aesthetic signals, leaves them without a mechanism to autonomously determine when to stop the composition process. To address these challenges, we introduce the Unified Aesthetic Foresight Model (UAFM), which seamlessly unifies both capabilities into a dual-head architecture over a shared backbone. Crucially, we formulate FOC as a deterministic Markov Decision Process and optimize a value head via a post-decision state temporal difference (TD) objective to recursively backpropagate sparse terminal aesthetic rewards to intermediate states. This enables the model to accurately estimate the potential of a partial outfit evolving into a compatible ensemble. Consequently, UAFM can compute step-wise marginal aesthetic gains for dynamic termination. Furthermore, this unified design strictly aligns aesthetic evaluation and complementary item prediction within a shared representation manifold, intrinsically regularizing the combinatorial search space. Extensive experiments on the Polyvore-Outfits dataset demonstrate that UAFM establishes a new state-of-the-art across both FOC and conventional fashion tasks. Extensive ablation studies further confirm that our post-decision state TD formulation provides the aesthetic value prediction necessary for autonomous dynamic termination, while validating the synergistic benefits of our unified dual-head architecture.