arXiv · 2609.24031
Video-STLayout Pre-training
Abstract
In recent years, pre-training has become fundamental to learning effective video representations, enabling strong transfer to downstream tasks. A popular framework in pre-training involves aligning features of a video encoder with that of another modality, for example, language or audio. We introduce Video-STLayout pre-training, a novel strategy for obtaining rich video representations informed by spatio-temporal layout of object bounding boxes. Object layouts can easily be obtained by applying an off-the-shelf object detector on the video frames. Our method uses a contrastive loss to align video features with the layout features from a trained layout encoder. We show the effectiveness of our approach in the task of activity recognition in complex scenes.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Akash Abdu Jyothi, Greg Mori. 2026-09-21. Video-STLayout Pre-training. https://arxiv.org/abs/2609.24031
Cite the original work for its findings. Save a collection to share your selection of sources.