Search arXiv⌕ Search

arXiv · 2609.30965

FRAM: Trajectory-Guided Visual Feature Selection for Compact Language-Conditioned Robot Manipulation

Abstract

Vision-Language-Action models achieve strong performance in robot manipulation, but often require large numbers of parameters. In this work, we propose the Future Representation Action Model (FRAM), a small policy that explicitly links the future end-effector trajectory to the current visual input. FRAM uses the image coordinates of the predicted trajectory as spatial pointers and reads local visual features related to the motion from the current image. This organizes the information for action generation into the reference position (Where), the visual state (What), and the future motion (Future). Trajectory labels are generated automatically from demonstrations and camera geometry, so no manual annotation is needed. With 138.7M parameters, including a frozen language encoder, FRAM reaches an average success rate of 92.2% over the four standard LIBERO suites, close to the 94.2% of $π_0$ with 3.3B parameters. Without extra training, it also reaches an average of 67.3% on LIBERO-Plus. Ablations confirm that both the future trajectory and the local visual features improve performance and robustness. On a real dual-arm UR5e, FRAM stacks cups using only wrist cameras, including choosing and switching between the left and right arms. These results show that selecting visual information based on future motion is an effective way to obtain both high performance and robustness in a small robot policy.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Hiroshi Ito, Hyogo Hiruma, Yoshiki Kanai, Takahiro Yoshida, Akira Kanazawa, Hiroki Yamada. 2026-09-25. FRAM: Trajectory-Guided Visual Feature Selection for Compact Language-Conditioned Robot Manipulation. https://arxiv.org/abs/2609.30965

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Structured-Diffuser: Diffusion with Task-Conditioned Structured Priors for Motion Planning

We propose Structured-Diffuser, a diffusion planner that embeds task and motion structure directly into the noise model. Unlike standard diffusion-based planners that rely on zero-mean, isotropic Gaussian corruption, we introduce task-conditioned structured Gaussians whose means and covariances are derived from Gaussian Process Motion Planning (GPMP), explicitly encoding trajectory smoothness and task semantics in the prior. We first formulate diffusion under a task-conditioned, non-isotropic Gaussian prior with closed-form forward and posterior expressions. Building on this formulation, our hierarchical design separates prior instantiation from trajectory denoising. At the upper level, sparse task-centric key states and timings are obtained, which instantiate a structured Gaussian prior (mean and covariance). At the lower level, the full trajectory is denoised under this prior, treating the upper-level outputs as noisy observations. Experiments across three motion-planning tasks show improved task success and training efficiency, with additional gains in data efficiency, trajectory smoothness, and position--velocity consistency where evaluated. Ablation studies further show that explicitly structuring the corruption process provides benefits beyond neurally conditioning the denoising network alone. Overall, our approach concentrates the prior's probability mass around task-relevant, temporally structured trajectories. We additionally demonstrate deployment on a physical G1 humanoid.

cs.RO↗

SANDO: Safe Autonomous Trajectory Planning for Dynamic Unknown Environments

This paper presents SANDO, a safe trajectory planner for 3D dynamic unknown environments. Existing soft-constraint planners are fast but do not guarantee collision-free paths, while hard-constraint methods typically ensure safety at the cost of longer computation. SANDO addresses this trade-off through three contributions. First, a heat map-based A* global planner steers the path away from high-risk regions, and a spatiotemporal safe flight corridor (STSFC) generator produces time-layered polytopes that inflate obstacles only by their worst-case reachable set at each time layer, rather than over the entire horizon. Second, trajectory optimization is formulated as a mixed-integer quadratic program with hard collision-avoidance constraints, and variable elimination reduces the number of decision variables. Third, a formal safety analysis establishes collision-free guarantees under explicit velocity-bound, size-bound, and estimation-error assumptions. Ablation studies confirm that variable elimination yields up to 7.4 times faster optimization and that STSFCs are critical for feasibility in dense dynamic environments. In simulations against state-of-the-art methods, SANDO achieves a 100% success rate across all forest and dynamic benchmark difficulty levels with no constraint violations, and perception-only experiments demonstrate the full perception-to-planning pipeline. Hardware experiments with fully onboard planning, perception, and localization demonstrate six safe flights in static environments and twelve among dynamic obstacles.

cs.RO↗

Cloak: Zero-Shot Cross-Embodiment Manipulation by Masking the End-Effector from the VLA

We present Cloak, a training recipe that endows a Vision-Language-Action (VLA) model with zero-shot cross-embodiment transfer by cloaking the end-effector from its own wrist camera. The end-effector occupies a large and consistent region of the wrist view and masking it allows for embodiment-agnostic visual reasoning. Cloak renders a mask in simulation from the robot's known geometry, accurately and in real time, with no segmentation or generative models. During training, we augment the mask so the model generalizes to embodiments unseen at training time. We demonstrate the recipe with Cloak-VLA, a VLA trained with Cloak on a single parallel-jaw gripper dataset. No data of new embodiments is ever collected. Cloak-VLA transfers zero-shot to various unseen embodiments, including another gripper, another arm, and a five-fingered hand, while preserving the source embodiment's performance. By decoupling the wrist view from its own embodiment, Cloak allows data to outlive the hardware it was collected on.

cs.RO↗