Search arXiv⌕ Search

arXiv subjects

Md. Safayet Islam

Publications and source records attributed to Md. Safayet Islam.

2 recordsLinked to original sources

Query-aligned video frame selection for long video understanding

Multimodal large language models (MLLMs) process multimodal inputs by converting text, images, and videos into token sequences that are subsequently processed by a backbone language model. While MLLMs have achieved excellent performance in understanding the content of individual images, video understanding remains significantly more difficult because videos contain large number of video frames. MLLMs typically process only a subset of these frames, usually ranging from 8 to 64. MLLMs usually sample frames uniformly, regardless of their relevance to the question being answered. To address this limitation, several training-free, model-agnostic methods for selecting question-relevant frames have recently been proposed. In this work, we introduce a frame-selection method designed specifically for multiple-choice questions. We extend the query text by appending semantic cues derived from the answer choices and employ a direct query-frame alignment scoring mechanism. To the best of our knowledge, our method is the first to directly utilize answer choices as inference-time cues for selecting frames relevant to answering a question. The method first constructs a compact candidate pool by subsampling video frames at a fixed rate. The frames are then scored according to their maximum cosine similarity across all question-answer pairs to identify the most relevant frames for a given query. This approach preserves a fixed token budget while improving the relevance of the visual evidence provided to the downstream MLLM. We evaluate the effectiveness of our frame-selection method on the MLVU, Video-MME, and LongVideoBench benchmarks using three MLLMs: LLaVA-Mini, Qwen2-VL, and LLaVA-Video. Experimental results demonstrate that answer-aware frame selection generally outperforms uniform sampling and existing training-free frame-selection methods under the same frame budget.

cs.CV↗

Evaluation of MLLM-Agnostic Plug-and-Play Keyframe Selection Methods for Long Video Understanding

Multimodal large language models (MLLMs) cannot process every frame of a long video because of limitations in visual-token and computational budgets. Three primary approaches have been proposed to enhance their long-video understanding capabilities: (i) Retraining an MLLM on a large video corpus and/or extending its input length; (ii) Training an adapter for a specific MLLM that takes the entire video and the query as input and selects the most relevant video frames; and (iii) Developing a training-free, plug-and-play (PaP) adapter that is MLLM-agnostic. We refer to the third approach as PaP keyframe selection. A PaP method may use only candidate video frames without considering the query, or it may use both candidate video frames and the query. The first approach is prohibitively expensive. The second approach requires substantial training time and computational resources, but it is accessible to many because an adapter contains significantly fewer trainable parameters than an entire MLLM. The third approach has the lowest computational cost and is therefore broadly accessible. To the best of our knowledge, only five PaP methods have been reported within the past year. All of these methods have been evaluated on one or more video question-answering benchmarks and have demonstrated improvements in long-video understanding. However, the methods were evaluated on different benchmarks using different MLLMs. We present a comprehensive evaluation of these five methods using three MLLMs across three long-video understanding benchmarks. Our results show that QAaF achieves the best performance in 13 of the 15 aggregate evaluation settings, while FOCUS ranks second overall. These results provide a common experimental reference for comparing training-free keyframe-selection methods for MLLMs.

cs.CV↗