arXiv · 2609.28159
Depth-Guided Contrastive Learning for 2D Representations with 3D Spatial Awareness
Abstract
Standard contrastive learning frameworks are mainly designed from a semantic perspective, yet learning 2D visual representations that preserve 3D spatial structure is also important for scene understanding. In this work, we propose Depth-Guided Contrastive Learning (DGCL), a simple auxiliary objective that injects 3D spatial awareness into 2D contrastive representation learning. Our key idea is to use depth to convert local 3D proximity into contrastive similarity: pixels that are closer in 3D space are encouraged to have more similar representations than pixels that are farther apart. Instead of relying on absolute depth values, DGCL formulates supervision through relative 3D distance comparisons among randomly sampled pixels, making the objective invariant to depth scale, efficient to compute, and easy to integrate into existing contrastive frameworks. Experiments across different datasets and models show that DGCL consistently improves 2D representation learning and benefits semantic downstream tasks by stronger spatial and geometric understanding. The code is available on https://github.com/LeungTsang/DGCL.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Liang Zeng, Maarten Vergauwen. 2026-09-23. Depth-Guided Contrastive Learning for 2D Representations with 3D Spatial Awareness. https://arxiv.org/abs/2609.28159
Cite the original work for its findings. Save a collection to share your selection of sources.