arXiv · 2609.01654
MELON: A Large-Scale Dataset for Multi-Event Text-to-Long-Video Retrieval
Abstract
Existing text-video retrieval datasets primarily consist of short-form clips containing a single dominant event. While suitable for measuring basic vision-language alignment, they are limited in capturing real-world retrieval scenarios, where long-form videos naturally contain multiple semantically distinct events and a single text query may correspond to several non-contiguous temporal segments. To bridge this gap, we introduce MELON, the first large-scale dataset designed to extend text-video retrieval to long-form videos featuring complex, multi-event structures. MELON explicitly annotates multiple event intervals per video along with their corresponding textual descriptions, enabling both training and evaluation of multi-event understanding in long, untrimmed videos. In addition, we propose a multi-event aware loss that encourages models to differentiate between full-event and partial-event matches, yielding substantial improvements in retrieval accuracy. Together, the MELON dataset and our proposed loss establish a robust foundation for expanding text-to-video retrieval to complex long-form scenarios and provide a more realistic evaluation setting for future research in the field.
Explore related subjects
Keep this discovery
Chan Hur, SeungWoo Song, Jeong-hun Hong, Won Jun Oh, Hyeyoung Park, KyungTae Lim. 2026-08-31. MELON: A Large-Scale Dataset for Multi-Event Text-to-Long-Video Retrieval. https://arxiv.org/abs/2609.01654
Cite the original work for its findings. Save a collection to share your selection of sources.
Discover connections
Connections use source metadata and explicit phrase matches, not verified experimental comparisons.