arXiv · 2603.09632
X-GS: An Extensible Framework for Perceiving and Thinking with 3D Gaussian Splatting
Abstract
3D Gaussian Splatting (3DGS) has emerged as a powerful technique for novel view synthesis, subsequently extending into numerous spatial AI applications. However, most existing 3DGS methods operate in isolation, focusing on specific domains. In this paper, we introduce X-GS, an extensible framework that integrates previously isolated 3DGS methods into the perception module of a VLM for spatial tasks, with two major components: the $\textit{Perceiver}$ and the $\textit{Thinker}$. The $\textit{Perceiver}$ performs online 3DGS-based SLAM with semantic distillation and outputs semantic Gaussians from unposed video streams. It leverages recent vision foundation models for stronger geometric priors, and we introduce three novel optimizations to improve semantic distillation efficiency. The $\textit{Thinker}$ interfaces diverse VLMs with these semantic Gaussians, unlocking spatial multimodal capabilities such as 3D visual grounding and scene captioning. Experimental results on diverse benchmarks demonstrate the efficiency and newly unlocked multimodal capabilities of the X-GS framework.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yueen Ma, Zenglin Xu, Irwin King. 2026-09-20. X-GS: An Extensible Framework for Perceiving and Thinking with 3D Gaussian Splatting. https://arxiv.org/abs/2603.09632
Cite the original work for its findings. Save a collection to share your selection of sources.