arXiv · 2610.05029
SPACE-CLIPv2: Decoding Local Geometry from Frozen CLIP for Monocular Depth Estimation
Abstract
Vision-language foundation models such as CLIP provide strong semantic representations, but their patch tokens are not directly optimized for dense metric geometry. SPACE-CLIP showed that frozen CLIP features can support monocular depth estimation through layer-group feature fusion, yet it leaves open how neighboring CLIP tokens should be combined to recover fine local structure. We present SPACE-CLIPv2, a frozen-backbone depth decoder that aggregates fixed local neighborhoods in CLIP token space. At selected decoder stages, the model samples a fixed token stencil, predicts aggregation weights, and injects the resulting response through a gated residual update. A token-space high-pass branch further preserves shallow local contrast. On NYU Depth V2, SPACE-CLIPv2 improves over a matched SPACE-CLIP baseline, while five-seed experiments consistently favor fixed over learned-offset sampling. Zero-shot iBims-1 evaluation further improves boundary and planar-geometry measures. These results support constrained local token aggregation as a practical mechanism for decoding geometry from frozen CLIP representations.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Hyun Song, Taewan Cho, Kangmin Kim, Andrew Jaeyong Choi. 2026-10-04. SPACE-CLIPv2: Decoding Local Geometry from Frozen CLIP for Monocular Depth Estimation. https://arxiv.org/abs/2610.05029
Cite the original work for its findings. Save a collection to share your selection of sources.