arXiv · 2607.04872
EventCoT: Event-centric Video Chain-of-thought for Reasoning Temporal Localization
Abstract
Reasoning temporal localization (RTL) requires a model to generate an answer that itself contains the time interval supporting it, coupling high-level reasoning with temporal grounding in a single response. To tackle this challenge, we propose the first event-centric video chain-of-thought framework, dubbed EventCoT. EventCoT first performs event-centric tokenization, converting the video into compact event tokens that enable efficient identification of question-relevant events. It then reasons within these events to generate the answer, grounding the time interval via embedding matching that aligns placeholder tokens with visual embeddings. EventCoT achieves state-of-the-art results on ActivityNet-RTL while using substantially fewer visual tokens than previous work, and attains strong zero-shot results on the grounded video question answering benchmark ReXTime. Our code will be released for research purposes.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Youngkil Song, Yoonjae Baek, Dongwon Kim, Inho Kim, Dongkeun Kim, Suha Kwak. 2026-07-06. EventCoT: Event-centric Video Chain-of-thought for Reasoning Temporal Localization. https://arxiv.org/abs/2607.04872
Cite the original work for its findings. Save a collection to share your selection of sources.