arXiv · 2609.25832
PartLLM: A Unified Multimodal Foundation for 3D Part Segmentation
Abstract
Part segmentation is a fundamental problem in computer graphics and 3D vision. Recent works have expanded 3D part segmentation beyond fixed taxonomies, but existing approaches typically only address a specific setting, such as text-guided part segmentation or point-based interaction. In this work, we argue that these settings can be unified as an intent-conditioned generative problem, where different prompts specify the desired part decomposition. To this end, we introduce PartLLM, a unified multimodal model that formulates 3D part segmentation as autoregressive semantic decomposition. Conditioned on an input shape and a user prompt, PartLLM autoregressively generates semantic part hypotheses as queries for mask prediction and feeds them to a decomposition-aware decoder that jointly predicts coherent part masks. This unified design supports text-guided part segmentation, interactive segmentation, and full-shape semantic decomposition with controllable granularity within a single model. Extensive experiments across these task settings show that PartLLM consistently outperforms task-specific baselines, demonstrating the effectiveness of unifying 3D part segmentation under an intent-conditioned generative formulation.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Zhe Zhu, Yiheng Zhang, Peng Li, Zixing Zhao, Honghua Chen, Yaqing Zhang, Le Wan, Zhiyang Dou, Cheng Lin, Yuan Liu, Mingqiang Wei, Wenping Wang. 2026-09-22. PartLLM: A Unified Multimodal Foundation for 3D Part Segmentation. https://arxiv.org/abs/2609.25832
Cite the original work for its findings. Save a collection to share your selection of sources.