Search arXiv⌕ Search

arXiv subjects

Madhav Kataria

Publications and source records attributed to Madhav Kataria.

3 recordsLinked to original sources

Re:Cognize -- Open-Set Comic Character Re-Identification

A manga reader meets a character on one page and knows them on sight a hundred pages later, without ever being handed a cast list. Re-identifying comic characters demands the same, open-set and sequential: pages arrive as a stream in reading order, new faces appear before anyone names them, and the cast is assembled as the story is read. $\textbf{Re:Cognize}$ evaluates recognition as the story is read, not against a cast handed over in advance: four protocols on one query stream, from closed-set retrieval to a cast the model must build and grow itself. The surprise is where models fail. Recognising is close to solved: one reference image per character already ranks as well as a gallery built in advance. Knowing what to believe is not: a model that adds its own matches makes its cast worse, while the same growth with correct labels would gain over twenty points of top-1 accuracy. The bottleneck is acceptance, not vision, and one comparison decides it: an addition pays exactly when it is right more often than the cast already was on the queries it takes over. The comparison has nothing to fit, and measured on half of a new corpus it calls the other half correctly. $\textbf{ReCast}$ puts it to work with nothing fitted on data: a cast sheet of one running average per character, grown only where the page itself vouches for a crop. It recovers a third to two thirds of what perfect labels would, depending on whether the cast starts from random examples or from first appearances. Re:Cognize measures whether a model can read along; ReCast is a cast that does. Our claims are on identity maintenance, recognising characters already met; the emergence of new ones is measured as a diagnostic under a fixed reference rule, and we propose no method for it.

cs.CV↗

PlotTwist: A Creative Plot Generation Framework with Small Language Models

Creative plot generation presents a fundamental challenge for language models: transforming a concise premise into a coherent narrative that sustains global coherence, character development, pacing, tone consistency, and emotional progression. Although recent Large Language Models (LLMs) demonstrate strong fluency on general-purpose tasks, they require preference alignment to perform well on domain-specific tasks such as creative plot generation. However, conducting such alignment at the scale of frontier LLMs is computationally prohibitive, significantly limiting accessibility and practical deployment. To address this, we present PlotTwist, a structured framework that enables Small Language Models (SLMs) with $\leq$3B active parameters to generate high-quality, premise-conditioned plots competitive with frontier systems of vastly greater parameter scale. Our approach decomposes generation into three specialized components: (1) an Aspect Rating Reward Model, trained via a novel Positive-Negative prompting strategy; (2) a Mixture-of-Experts (MoE) plot generator aligned via Direct Preference Optimization (DPO); and (3) an Agentic Evaluation module using a cross-family jury for unbiased, independent post-hoc assessment. Extensive experiments demonstrate that PlotTwist consistently outperforms all baselines, including frontier models, across multiple Narrative Quality Dimensions (NQDs), achieving higher win rates against every baseline except the strongest, with which it remains competitive. Further validation confirms strong sensitivity to narrative quality, as the framework reliably distinguishes plots derived from critically acclaimed versus widely panned screenplays. Together, these results establish structured, preference-based alignment as a resource-efficient approach to high-quality creative plot generation. Project page: https://abhinavthorat.github.io/plottwist/

cs.CL↗

Re:Verse -- Can Your VLM Read a Manga?

Current Vision Language Models (VLMs) demonstrate a critical gap between surface-level recognition and deep narrative reasoning when processing sequential visual storytelling. Through a comprehensive investigation of manga narrative understanding, we reveal that while recent large multimodal models excel at individual panel interpretation, they systematically fail at temporal causality and cross-panel cohesion, core requirements for coherent story comprehension. We introduce a novel evaluation framework that combines fine-grained multimodal annotation, cross-modal embedding analysis, and retrieval-augmented assessment to systematically characterize these limitations. Our methodology includes (i) a rigorous annotation protocol linking visual elements to narrative structure through aligned light novel text, (ii) comprehensive evaluation across multiple reasoning paradigms, including direct inference and retrieval-augmented generation, and (iii) cross-modal similarity analysis revealing fundamental misalignments in current VLMs' joint representations. Applying this framework to Re:Zero manga across 11 chapters with 308 annotated panels, we conduct the first systematic study of long-form narrative understanding in VLMs through three core evaluation axes: generative storytelling, contextual dialogue grounding, and temporal reasoning. Our findings demonstrate that current models lack genuine story-level intelligence, struggling particularly with non-linear narratives, character consistency, and causal inference across extended sequences. This work establishes both the foundation and practical methodology for evaluating narrative intelligence, while providing actionable insights into the capability of deep sequential understanding of Discrete Visual Narratives beyond basic recognition in Multimodal Models. Project Page: https://re-verse.vercel.app

cs.CV↗