arXiv · 2609.40222
LOCI: Spatial Linear Memory for Streaming World Models
Abstract
When a camera revisits a previously observed region, a video world model should reproduce what was there before. This requires both remembering past observations and retrieving the right one for the current viewpoint. Key-value caches preserve visual detail but grow with video length; recurrent memory is compact but compresses history into a fixed-size state, so individual past observations are no longer directly accessible. We introduce LOCI, a hybrid spatial-memory architecture that keeps both representations. In half of the transformer blocks, main attention keeps a key-value cache of past observations; in the other half, it is restricted to the current chunk and complemented by a recurrent linear-attention memory whose reads and writes are conditioned on projective camera geometry, so viewpoint enters both memory addressing and stored content. Recurrent readouts flow into subsequent cache-backed blocks and supply their queries with accumulated scene context. On the public MIND memory benchmark and on held-out recorded trajectories, LOCI reproduces revisited content more faithfully than representative world models and a same-recipe full-softmax model; with full history, it lowers peak memory at equal length by about 30% relative to full softmax. With a bounded bank of retained observations, it streams long videos at constant memory and remains more faithful than full softmax under the same budget.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Ji Xia, Tingting Liao, Xuezhi Liang, Hao Li, Guangyi Liu. 2026-09-30. LOCI: Spatial Linear Memory for Streaming World Models. https://arxiv.org/abs/2609.40222
Cite the original work for its findings. Save a collection to share your selection of sources.