What Should We Ask Next? Retrieval-Aware Question Learning for Interactive ReID
Interactive retrieval with partial evidence constitutes a sequential information-acquisition problem: an agent must choose questions that acquire useful evidence for the next retrieval update. Existing systems train this decision by imitating an offline ordering of candidate QA pairs. However, a question's value depends on the response it elicits and its downstream effect on retrieval. We establish that candidate discriminativeness and perceived usefulness provide weak supervision for this objective, then introduce RAVEL, a retrieval-aware online reinforcement learning framework for interactive person re-identification. RAVEL initializes from supervised question generation, observes the current Top-4 candidates directly, and optimizes the question policy with rank feedback from the full question-answer-retrieval loop. Experiments on Interactive-PEDES show that RAVEL delivers progressively stronger retrieval performance across five interaction rounds. Further analysis shows that RAVEL allocates more of its interaction budget to localized open-ended prompts targeting specific attributes.