arXiv · 2609.26118
GDLAM: Group-Disentangled Latent Action Model for Highly Disentangled Embodied Pretraining
Abstract
Latent action models (LAMs) learn action-related representations from action-free videos via self-supervised future prediction, offering a scalable paradigm for embodied intelligence pretraining. However, existing LAMs collapse heterogeneous sources of visual change, including camera motion, object dynamics, and interaction events, into a single latent vector, resulting in entangled representations with limited semantic structure and consequently restricting world model controllability and VLA policy generalization. We introduce the Group-Disentangled Latent Action Model (GDLAM), a latent action model whose code is factorized by construction into N groups, each with an independent variational bottleneck and a spatially gated routing pathway, and trained with a set of information-geometric objectives: mutual exclusivity, group and gate sparsity, and static-dynamic orthogonality, that make the groups mutually causally distinct rather than merely decorrelated. Quantitatively, intervening on any single group changes only that group and leaves the others intact, and GDLAM improves label-free disentanglement metrics, including Modularity, MIG, and DCI, by wide margins over a strong unstructured LAM. Notably, this factorization is not at the expense of action information: across three mutual-information estimators and a linear probe, the grouped code is more informative than monolithic baselines both in- and out-of-distribution. As supporting evidence that the disentangled code is a reusable pretraining currency, we further transfer it to two downstream regimes: (1) World Modeling: World models pretrained with GDLAM achieve superior rollout fidelity and action-following capability compared with SOTA baselines. (2) VLA Policies: Pretraining with GDLAM substantially improves task success rates over previous methods across multiple simulation benchmarks and real-world robotic manipulation tasks
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Jiarui Yang, Jiawei Li, Jiale Zhang, Hang Guo, Wen Huang, Maowei Hu, Tao Dai, Shu-Tao Xia. 2026-08-10. GDLAM: Group-Disentangled Latent Action Model for Highly Disentangled Embodied Pretraining. https://arxiv.org/abs/2609.26118
Cite the original work for its findings. Save a collection to share your selection of sources.