arXiv · 2609.36755
Drag as Evidence: Motion-Grounded Latent Recomposition for Drag-Based Editing
Abstract
Modern image editors excel at semantic manipulation and visual synthesis, yet remain limited in precise spatial control, motivating the development of drag-based editing. However, existing drag-based methods often struggle to balance drag accuracy with natural, plausible, and intent-aligned generation. We propose MoRe-Drag, a motion-grounded drag-based editing method. Our key insight is to treat pixel-space warping as coarse motion evidence, and to inject this evidence into the generative sampling trajectory. Specifically, MoRe-Drag performs region-aware latent recomposition over refinement, inpainting, and anchor regions, coupled with stage-adaptive conditioning that progressively shifts from motion-grounded structure formation to semantic refinement. We further support an instruction-free interface by adapting the MLLM-based text encoder for drag-aware instruction inference. Experiments on DragBench-SR and DragBench-DR show that MoRe-Drag substantially improves drag precision over strong base editors and achieves superior drag accuracy among SOTA drag-based methods, while delivering strong semantic consistency and visually realistic results. Code and dataset will be publicly released.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Xinyu Pu, Hongsong Wang, Jie Gui, Pan Zhou. 2026-09-29. Drag as Evidence: Motion-Grounded Latent Recomposition for Drag-Based Editing. https://arxiv.org/abs/2609.36755
Cite the original work for its findings. Save a collection to share your selection of sources.