Search arXiv⌕ Search

arXiv subjects

Dewang Sultania

Publications and source records attributed to Dewang Sultania.

6 recordsLinked to original sources

HeroFrame-Bench: Reference-Anchored Evaluation via Rubric--Ranking Co-Evolution for Movie Hero Frame Selection

Hero frames are in-film stills used as source imagery for theatrical posters, streaming cover art, film database listings, and other promotional placements. As the first visual entry point, they shape audiences' initial impressions of the movie and their subsequent willingness to watch it. Selecting these frames, a task we term hero frame selection, requires balancing content relevance with aesthetic appeal. A related task is keyframe selection, yet its benchmarks prioritize relevance over aesthetics, using either finite annotations that exclude valid alternatives or VideoQA that entangles selection quality with downstream model capability. We therefore introduce HeroFrame-Bench, built through a scalable VLM-as-a-Judge framework. We construct multimodal contexts from diverse metadata to ground a VLM judge that scores selected frames using our Reference-anchored Percentile. The percentile is obtained by inserting each frame into reusable, pre-ranked reference chains, enabling direct, extensible, and reliable evaluation. To reduce ambiguity and improve consistency in these subjective judgements, we further propose Rubric-Ranking Co-Evolution, which generates movie-specific rubrics to condition the VLM judge and refines rubrics jointly with the resulting rankings. Within this process, we introduce several verifiable signals, most notably the Inverted Rubric Attack, to select robust rubrics. Finally, HeroFrame-Bench is instantiated over 204 movies with 2,031 reference chains and 1,970 learned rubrics. We build an annotation interface for human-alignment studies which show that our construction design improves VLM agreement with human from 77.56% to 83.78%. Evaluation on multiple methods show that hero frame selection remains challenging.

cs.CV↗

CueKFS: Agentic Cue-Driven Keyframe Selection for Long Video Understanding

Keyframe selection (KFS) has long produced compact video summaries for browsing and retrieval, and representative frames for thumbnails. More recently, when conditioned on a question, KFS provides an alternative to uniform sampling for long-video question answering by selecting frames that are more relevant to the question. Most methods rank frames by similarity to the question. Yet a relevant frame may score poorly when the question combines subjects or moments that no single frame shows, or requires implicit information absent from its wording. Other methods try to break down the question into subqueries, but suffer from inaccurate decomposition due to their static initial context. Therefore, we propose CueKFS, a training-free method that reformulates question--frame matching as comparing frames against a set of dynamically generated visual cues. From an initial set of salient frames, we decompose the question into cues. Each cue concurrently probes the video to navigate to its own evidence. A reasoning VLM then agentically revises the cue set against its evidence to re-explore the video. CueKFS then allocates the budget across the surviving cues. Across three benchmarks, CueKFS establishes state-of-the-art results in all 27 evaluated settings with available prior results, achieving budget-averaged gains of up to +4.54% over the previous baseline and a median of only two VLM calls. We further provide a detailed behavioral analysis of CueKFS, showing that agentic cue refinement drives active re-exploration of the video, yielding relative similarity gains of up to 92% over the initial context.

cs.CV↗

FUSE : Failure-aware Usage of Subagent Evidence for MultiModal Search and Recommendation

Multimodal creative assistants decompose user goals and route tasks to subagents for layout, styling, retrieval, and generation. Retrieval quality is pivotal, yet failures can arise at several stages: understanding user intent, choosing content types, finding candidates (recall), or ranking results. Meanwhile, sending and processing images is costly, making naive multimodal approaches impractical. We present FUSE: Failure-aware Usage of Subagent Evidence for MultiModal Search and Recommendation. FUSE replaces most raw-image prompting with a compact Grounded Design Representation (GDR): a selection aware JSON of canvas elements (image, text, shape, icon, video, logo), structure, styles, salient colors, and user selection provided by the Planner team. FUSE implements seven context budgeting strategies: comprehensive baseline prompting, context compression, chain-of-thought reasoning, mini-shot optimization, retrieval-augmented context, two-stage processing, and zero-shot minimalism. Finally, a pipeline attribution layer monitors system performance by converting subagent signals into simple checks: intent alignment, content-type/routing sanity, recall health (e.g., zero-hit and top-match strength), and ranking displacement analysis. We evaluate the seven context budgeting variants across 788 evaluation queries from diverse users and design templates (refer Figure 3). Our systematic evaluation reveals that Context Compression achieves optimal performance across all pipeline stages, with 93.3% intent accuracy, 86.8% routing success(with fallbacks), 99.4% recall, and 88.5% NDCG@5. This approach demonstrates that strategic context summarization outperforms both comprehensive and minimal contextualization strategies.

cs.IR↗

RouteNator: A Router-Based Multi-Modal Architecture for Generating Synthetic Training Data for Function Calling LLMs

This paper addresses fine-tuning Large Language Models (LLMs) for function calling tasks when real user interaction data is unavailable. In digital content creation tools, where users express their needs through natural language queries that must be mapped to API calls, the lack of real-world task-specific data and privacy constraints for training on it necessitate synthetic data generation. Existing approaches to synthetic data generation fall short in diversity and complexity, failing to replicate real-world data distributions and leading to suboptimal performance after LLM fine-tuning. We present a novel router-based architecture that leverages domain resources like content metadata and structured knowledge graphs, along with text-to-text and vision-to-text language models to generate high-quality synthetic training data. Our architecture's flexible routing mechanism enables synthetic data generation that matches observed real-world distributions, addressing a fundamental limitation of traditional approaches. Evaluation on a comprehensive set of real user queries demonstrates significant improvements in both function classification accuracy and API parameter selection. Models fine-tuned with our synthetic data consistently outperform traditional approaches, establishing new benchmarks for function calling tasks.

cs.LG↗

Domain-specific Question Answering with Hybrid Search

Domain specific question answering is an evolving field that requires specialized solutions to address unique challenges. In this paper, we show that a hybrid approach combining a fine-tuned dense retriever with keyword based sparse search methods significantly enhances performance. Our system leverages a linear combination of relevance signals, including cosine similarity from dense retrieval, BM25 scores, and URL host matching, each with tunable boost parameters. Experimental results indicate that this hybrid method outperforms our single-retriever system, achieving improved accuracy while maintaining robust contextual grounding. These findings suggest that integrating multiple retrieval methodologies with weighted scoring effectively addresses the complexities of domain specific question answering in enterprise settings.

cs.CL↗

Retrieval Augmented Generation for Domain-specific Question Answering

Question answering (QA) has become an important application in the advanced development of large language models. General pre-trained large language models for question-answering are not trained to properly understand the knowledge or terminology for a specific domain, such as finance, healthcare, education, and customer service for a product. To better cater to domain-specific understanding, we build an in-house question-answering system for Adobe products. We propose a novel framework to compile a large question-answer database and develop the approach for retrieval-aware finetuning of a Large Language model. We showcase that fine-tuning the retriever leads to major improvements in the final generation. Our overall approach reduces hallucinations during generation while keeping in context the latest retrieval information for contextual grounding.

cs.CL↗