arXiv · 2609.37488
FORUM: Frozen Outputs Reconciled Using Model Agreement for Visual Grounding
Abstract
Frozen multimodal large language models (MLLMs) now solve standard referring expression comprehension with a single prompted call, yet on adversarial benchmarks with same-category distractors and negation, even the largest models are confidently wrong, and resampling repeats the error. Models built from different data and architectures rarely fall for the same confounder, so their agreement is a strong label-free signal of the correct target. We present FORUM, a training-free test-time fusion of frozen MLLMs guided by two fixed geometric rules: agreement-based selection keeps the region supported by the most distinct models, and medoid localization returns an actual member box instead of a coordinate average, so one loose prediction cannot shift the answer. Fusing three open MLLMs, FORUM surpasses the 397B-parameter published reference by a relative 5% in mean accuracy on the adversarial Ref-Adv-s benchmark, and a plain averaging ensemble by 15%. The gains transfer to standard RefCOCO+, and a balanced lineup with no dominant member still surpasses the 397B model by 5%.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Taiyo Sato, Takamasa Sanda, Keisuke Maeda, Takahiro Ogawa, Miki Haseyama, Shunya Nagashima. 2026-09-27. FORUM: Frozen Outputs Reconciled Using Model Agreement for Visual Grounding. https://arxiv.org/abs/2609.37488
Cite the original work for its findings. Save a collection to share your selection of sources.