Search arXiv⌕ Search

arXiv · 2609.30436

WALT: Learning World-Model-Aligned Latent Trajectories for Autonomous Driving

Abstract

Driving world models learn rich predictive representations of the surrounding environment from visual observations, yet accurate visual prediction does not necessarily translate into effective trajectory planning. We argue that a key bottleneck lies in the mismatch between visual world states and raw geometric trajectories, which may limit the planner's ability to exploit action-relevant semantics encoded by the world model. To address this issue, we propose World-Model Alignment for Latent Trajectories (WALT), which learns a compact generative trajectory latent space by transferring information from a frozen pretrained driving world model without modifying the world model itself. Rather than directly generating raw waypoints, WALT maps them into compact representations through a dual-branch trajectory autoencoder and transfers semantic knowledge from the frozen visual world model into this trajectory space, encouraging the learned action representation to capture scene-level cues relevant to future motion and planning. Beyond our proposed formulation, we systematically study latent learning based on Joint-Embedding Predictive Architectures (JEPA) and feature alignment following Representation Alignment (REPA) to investigate how trajectory-only representation learning affects downstream planning. We evaluate WALT on the NAVSIM benchmarks. Relative to the raw-waypoint baseline, WALT improves PDMS from 89.4 to 89.8 on NAVSIMv1 and EPDMS from 87.3 to 87.9 on NAVSIMv2 while reducing trajectory planner FLOPs by 30.5%. These results suggest that preserving world representations while extracting action-relevant information provides an effective interface for world-model-based trajectory planning.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Mingkai Jia, Jiaxin Guo, Zhijian Shu, Jiawei Xu, Mingxiao Li, Jintao Cheng, Ping Tan, Wei Yin. 2026-09-24. WALT: Learning World-Model-Aligned Latent Trajectories for Autonomous Driving. https://arxiv.org/abs/2609.30436

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Structured-Diffuser: Diffusion with Task-Conditioned Structured Priors for Motion Planning

We propose Structured-Diffuser, a diffusion planner that embeds task and motion structure directly into the noise model. Unlike standard diffusion-based planners that rely on zero-mean, isotropic Gaussian corruption, we introduce task-conditioned structured Gaussians whose means and covariances are derived from Gaussian Process Motion Planning (GPMP), explicitly encoding trajectory smoothness and task semantics in the prior. We first formulate diffusion under a task-conditioned, non-isotropic Gaussian prior with closed-form forward and posterior expressions. Building on this formulation, our hierarchical design separates prior instantiation from trajectory denoising. At the upper level, sparse task-centric key states and timings are obtained, which instantiate a structured Gaussian prior (mean and covariance). At the lower level, the full trajectory is denoised under this prior, treating the upper-level outputs as noisy observations. Experiments across three motion-planning tasks show improved task success and training efficiency, with additional gains in data efficiency, trajectory smoothness, and position--velocity consistency where evaluated. Ablation studies further show that explicitly structuring the corruption process provides benefits beyond neurally conditioning the denoising network alone. Overall, our approach concentrates the prior's probability mass around task-relevant, temporally structured trajectories. We additionally demonstrate deployment on a physical G1 humanoid.

cs.RO↗

SANDO: Safe Autonomous Trajectory Planning for Dynamic Unknown Environments

This paper presents SANDO, a safe trajectory planner for 3D dynamic unknown environments. Existing soft-constraint planners are fast but do not guarantee collision-free paths, while hard-constraint methods typically ensure safety at the cost of longer computation. SANDO addresses this trade-off through three contributions. First, a heat map-based A* global planner steers the path away from high-risk regions, and a spatiotemporal safe flight corridor (STSFC) generator produces time-layered polytopes that inflate obstacles only by their worst-case reachable set at each time layer, rather than over the entire horizon. Second, trajectory optimization is formulated as a mixed-integer quadratic program with hard collision-avoidance constraints, and variable elimination reduces the number of decision variables. Third, a formal safety analysis establishes collision-free guarantees under explicit velocity-bound, size-bound, and estimation-error assumptions. Ablation studies confirm that variable elimination yields up to 7.4 times faster optimization and that STSFCs are critical for feasibility in dense dynamic environments. In simulations against state-of-the-art methods, SANDO achieves a 100% success rate across all forest and dynamic benchmark difficulty levels with no constraint violations, and perception-only experiments demonstrate the full perception-to-planning pipeline. Hardware experiments with fully onboard planning, perception, and localization demonstrate six safe flights in static environments and twelve among dynamic obstacles.

cs.RO↗

Cloak: Zero-Shot Cross-Embodiment Manipulation by Masking the End-Effector from the VLA

We present Cloak, a training recipe that endows a Vision-Language-Action (VLA) model with zero-shot cross-embodiment transfer by cloaking the end-effector from its own wrist camera. The end-effector occupies a large and consistent region of the wrist view and masking it allows for embodiment-agnostic visual reasoning. Cloak renders a mask in simulation from the robot's known geometry, accurately and in real time, with no segmentation or generative models. During training, we augment the mask so the model generalizes to embodiments unseen at training time. We demonstrate the recipe with Cloak-VLA, a VLA trained with Cloak on a single parallel-jaw gripper dataset. No data of new embodiments is ever collected. Cloak-VLA transfers zero-shot to various unseen embodiments, including another gripper, another arm, and a five-fingered hand, while preserving the source embodiment's performance. By decoupling the wrist view from its own embodiment, Cloak allows data to outlive the hardware it was collected on.

cs.RO↗