Search arXivSearch

arXiv · 2307.15983

Instance-Wise Adaptive Tuning and Caching for Vision-Language Models

Abstract

Large-scale vision-language models (LVLMs) pretrained on massive image-text pairs have achieved remarkable success in visual representations. However, existing paradigms to transfer LVLMs to downstream tasks encounter two primary challenges. Firstly, the text features remain fixed after being calculated and cannot be adjusted according to image features, which decreases the model's adaptability. Secondly, the model's output solely depends on the similarity between the text and image features, leading to excessive reliance on LVLMs. To address these two challenges, we introduce a novel two-branch model named the Instance-Wise Adaptive Tuning and Caching (ATC). Specifically, one branch implements our proposed ConditionNet, which guides image features to form an adaptive textual cache that adjusts based on image features, achieving instance-wise inference and improving the model's adaptability. The other branch introduces the similarities between images and incorporates a learnable visual cache, designed to decouple new and previous knowledge, allowing the model to acquire new knowledge while preserving prior knowledge. The model's output is jointly determined by the two branches, thus overcoming the limitations of existing methods that rely solely on LVLMs. Additionally, our method requires limited computing resources to tune parameters, yet outperforms existing methods on 11 benchmark datasets.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Chunjin Yang, Fanman Meng, Shuai Chen, Mingyu Liu, Runtong Zhang. 2023-07-29. Instance-Wise Adaptive Tuning and Caching for Vision-Language Models. https://arxiv.org/abs/2307.15983

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Amanous: Distribution-Switching for Superhuman Piano Density on Disklavier

A player piano can strike more keys, across more of the register, and faster than any pianist can reach. Three traditions dominate composition in that region, namely Nancarrow's tempo canons, Xenakis' stochastic distributions, and L-system grammars. They have developed in isolation, and none of them accounts for the instrument itself. A Disklavier does not answer instantly, because a loud note reaches the string sooner than a soft one, so music written as if the mechanism were transparent arrives distorted. We present Amanous, a hardware-aware composition system for the Yamaha Disklavier that unifies the three traditions through distribution-switching, in which each grammar symbol selects an entire distributional regime rather than adjusting parameters within a fixed one. A four-layer pipeline carries a symbol from grammar to actuation-ready MIDI. An L-system fixes the macro-form, each symbol is mapped to distributions and a tempo-canon ratio, events are sampled and time-scaled, and a hardware layer pre-compensates velocity-dependent latency and enforces the key-reset time. Convergence points close a feedback loop, letting the music's own temporal structure trigger the next switch. Sections from different symbols remain statistically separable at the output, and ablating the L-system, the tempo canon, or the hardware compensation each degrades a distinct property. Under the modelled latency curve, pre-compensation takes the mean onset error from roughly 18 ms to below the instrument's 1 ms scanning resolution. Structured and random textures are most separable near 25 notes/s, and melodic metrics lose their discriminative power over the 40-100 notes/s band, above which the difference is consistent with distributional content rather than melodic order. All results are computational, nothing was played or measured on an instrument, and a psychoacoustic protocol is proposed for future work.

cs.MM

If You Hear It, Help Find It: Orthogonal Knowledge Distillation for Open-Vocabulary Audio-Visual Event Localization

Open-vocabulary audio-visual event localization (OV-AVEL) grounds a text-queried event in time from video, audio, and language. The supervision sources available to this task can differ in temporal-boundary reliability: on OV-AVEBench, our configured visual teacher gives more reliable boundary cues than the configured audio teacher, although the latter is a strong pretrained audio model and remains semantically informative. This is a setting-specific diagnostic rather than a universal ranking of vision and audio. We formulate the resulting challenge as supervision placement: which teacher signals may shape the localization decision, and which should remain auxiliary. Based on this view, we propose OV-OrthKD, a reliability-aware asymmetric distillation framework. Visual feature transfer shapes a decision-aligned representation, audio feature transfer enriches a complementary auxiliary subspace, a text prototype anchors seen/unseen category semantics, and an orthogonality loss limits directional overlap between the two teacher-specific projections. The student continues to use both modalities through query-aware fusion at inference, while the default training recipe keeps audio-teacher supervision off the segment-logit path. On OV-AVEBench, OV-OrthKD achieves 0.816 segment AP and improves F1@0.5 over the official fine-tuning baseline by 2.7 points overall and 3.4 points on unseen categories. Path-assignment, role-swap, corruption, and transfer analyses consistently support supervision placement as a task-specific design axis for OV-AVEL.

cs.MM

Adaptive Hierarchical Representation Alliance for Multimodal Learning

Multimodal models often align language, vision, and audio in a single final-layer latent space, implicitly assuming that task-relevant evidence emerges at the same semantic depth across modalities. Using layer-wise CKA analysis, we observe that this assumption leads to semantic granularity mismatch: textual cues usually require deeper contextual abstraction, whereas visual and acoustic cues often provide discriminative perceptual evidence in shallow or middle layers. This mismatch can flatten fine-grained modality-private cues and reduce reliability under noisy, imbalanced, or missing inputs. To address this, we proposed Adaptive Hierarchical Representation Alliance (AHRA), a hierarchical shared--private expert framework. AHRA factorizes each modality into shared and private streams across semantic levels, regularizes them with shared alignment and private decorrelation, routes shared information through a cross-modal expert, and enhances task-relevant private tokens with modality-specific experts guided by a sparsity-controlled soft-gating mechanism (foreground exam). A hierarchical co-fusion module then performs intra-level expert coordination and inter-level semantic selection. Experiments on six benchmarks across image-text classification, multimodal intent recognition, and trimodal sentiment analysis show that AHRA consistently improves over strong baselines and remains robust under noisy and missing-modality settings.

cs.MM