Toward Open-World Video Segmentation over Long Horizons
Long videos challenge open-world segmentation systems to continually discover objects and preserve their identities through disappearance and reappearance. Savvy, a zero-shot, semi-online, class-agnostic system for persistent object discovery and identity maintenance, and OGA, an evaluation suite that credits coherent part-level predictions when their granularity differs from the reference annotations. Savvy combines modular mask discovery, deferred admission based on accumulated evidence, and track consolidation to maintain an evolving object set. OGA allows multiple predictions to support one reference while preserving prediction IDs, pairing granularity-aware fidelity with diagnostics of identity bleeding and fragmented support. Across 142 ScanNet videos and 106 HM3D trajectories, evaluated using adaptively sampled prefixes capped at 1,500 source frames, Savvy outperforms DEVA+SAM and EntitySAM in VPQ_inf, STQ, and AQ under both conventional and OGA evaluation. Conventional VPQ_inf reaches 18.44 versus DEVA+SAM's 11.52 on ScanNet and 17.31 versus 6.58 on HM3D. VIPSeg comparisons reveal conventional one-to-one VPQ's sensitivity to annotation granularity. Controlled identity-severing and flicker tests show that OGA detects temporal failures even when frame-level masks remain unchanged. Together, these results demonstrate complementary advances: Savvy improves long-horizon segmentation and association, and OGA distinguishes coherent part-level support from temporal identity failures.