arXiv · 2609.34942
Resolution as a First-Class Decision: Task-Conditioned Routing for Efficient Multimodal Large Language Models
Abstract
The inference efficiency of Multimodal Large Language Models (MLLMs) is severely constrained by massive visual token sequences induced by high-resolution inputs, with computational cost scaling quadratically. Existing approaches primarily focus on downstream token compression, while overlooking a fundamental upstream inefficiency: input resolution is treated as a static, task-agnostic hyperparameter. We propose Task-Conditioned Resolution Routing (TCRR), which formulates visual compression as a task-conditioned decision and employs a lightweight cross-modal router that conditions backbone visual representations on textual semantics via feature-wise modulation and cross-attention to predict the minimal sufficient compression level per query. To support this, we curate a dataset of 500k samples across 12 task categories, labeled via a teacher-oracle pipeline to approximate Pareto-optimal compression scales. Extensive experiments across diverse architectures show that TCRR achieves a superior efficiency frontier, specifically reducing visual FLOPs by 40.9% and latency by 53.7% on Qwen3-VL-8B while preserving competitive performance. Further analysis of scaling behavior confirms that dynamically routing visual compression enables optimal resource allocation without modifying the MLLM backbone.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Zhiqiang Xia, Yang Li, Xinyuan Zhang, Yuchen Liu, Haoyu Lu, Jiaming Xu, Runyu Shi, Ying Huang. 2026-09-28. Resolution as a First-Class Decision: Task-Conditioned Routing for Efficient Multimodal Large Language Models. https://arxiv.org/abs/2609.34942
Cite the original work for its findings. Save a collection to share your selection of sources.