arXiv · 2609.36946
Optimizing VLP-aligned Multimodal Intent Representation with Correct Visual Instantiation for Zero-Shot Composed Image Retrieval
Abstract
ZS-CIR aims to retrieve a target image from a reference image and a modification text without paired supervision, typically by encoding composed queries as text-dominant representations within the image-text matching space of VLPs. However, queries reconstructed by visual pseudo-word learning or MLLM-based target reasoning often deviate from the native VLP representation space due to reference noise and coarse text fusion in the former, and verbose, weakly visually grounded descriptions in the latter. In this paper, we propose a unified ZS-CIR framework (named VMIR-CVI) to reconstruct multimodal composite queries from two complementary perspectives for optimizing VLP-compatible multimodal intent representation. First, it reasons and converts the multimodal intent into a unified textual description, aligning with the native text space of the VLP backbones to produce more retrieval-compatible textual queries. Second, it reconstructs the query representation with correctly decoupled visual instance cues, reducing reference noise while preserving target-relevant content. Specifically, a VLP-aligned Multimodal Intent Reasoning (VMIR) module injects few-shot VLP-style exemplars into chain-of-thought prompts, guiding the MLLM to generate target-consistent intent queries. A Training-free Visual Instance Disentanglement (TVID) module decouples fine-grained visual instances from global reference features without additional optimization. Finally, a lightweight Hybrid-modal Intent Alignment and Fusion (HIAF) module integrates the reasoned textual intent and disentangled visual cues into a unified hybrid-modal representation for robust ZS-CIR. Extensive experiments on three CIR benchmarks, namely CIRR, CIRCO and FashionIQ, show that VMIR-CVI significantly outperforms existing baselines and achieves new state-of-the-art performance. Code and trained models will be publicly released.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Xuri Ge, Chunhao Wang, Junchen Fu, Haokun Wen, Zhiwei Xu, Ying Zhou, Zhumin Chen, Pengjie Ren, Zhaochun Ren, Xin Xin. 2026-09-29. Optimizing VLP-aligned Multimodal Intent Representation with Correct Visual Instantiation for Zero-Shot Composed Image Retrieval. https://arxiv.org/abs/2609.36946
Cite the original work for its findings. Save a collection to share your selection of sources.