arXiv · 2609.29029
Exploiting answer-invariant redundancies in satellite imagery for efficient VLM inference on edge
Abstract
Onboard vision-language models could enable satellites to answer queries directly, but exhaustive tiled inference over high-resolution imagery is slow and energy-intensive. We identify answer-invariant token redundancy (AITR): image tiles and vision tokens that can be removed without changing the final answer. We present Rift, a two-stage system that performs query-conditioned tile pruning followed by elastic prefill to reduce token budget. We evaluate it on LLaVA-1.5 7B running on Jetson AGX Orin. Compared with exhaustive tiled inference, Rift reduces energy by 78% and latency by 69%, while increasing accuracy from 45% to 73%.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Ishani Janveja, Davis Zhang, Seoyul Oh, Deepak Vasisht. 2026-09-24. Exploiting answer-invariant redundancies in satellite imagery for efficient VLM inference on edge. https://arxiv.org/abs/2609.29029
Cite the original work for its findings. Save a collection to share your selection of sources.