arXiv · 2610.06021
Scalable Minimal-Change Learning for Controllable Image Editing
Abstract
Image editing should change only the attributes specified by an instruction while preserving everything else, yet current methods often make unintended changes. We treat this minimal-change principle as an optimization objective for instruction-based editing. Latent L1 regularization is a poor proxy for output locality in modern nonlinear generators and often requires supervision unavailable at scale. We instead optimize edit outcomes with reinforcement learning. An agentic vision-language reward model audits each source image, instruction, and edited image for two failure types: unimplemented requested changes and unintended changes. A group-level rubric merges and verifies these issues to provide consistent rewards across candidate edits without per-instruction human annotations. On FLUX.1 Kontext-dev, ARRO raises average EditScore from 5.21 to 5.88 across MinEval, MagicBrush, AnyBench, and Emu-Edit. On 600 evaluation examples, it reduces off-target pixel change by 8.4% relative to the base editor. Reward and SFT controls, blinded human evaluations, and transfer to OmniGen2 provide complementary evidence. Code: https://github.com/Showwwwwwwww/ARRO
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Shuo Chen, Fengming Huang, Yu Yao, Mingming Gong, Tongliang Liu. 2026-10-05. Scalable Minimal-Change Learning for Controllable Image Editing. https://arxiv.org/abs/2610.06021
Cite the original work for its findings. Save a collection to share your selection of sources.