Search arXiv⌕ Search

arXiv · 2610.02288

AdaptViT: Runtime-Adaptive Vision Transformer Deployment on Custom RISC-V

Abstract

Deploying Vision Transformers (ViTs) on low-power edge devices is challenging due to high computational demands. Conventional pruning frameworks require a separate compiled binary for each sparsity level, increasing storage overhead and limiting runtime adaptability. This paper presents an end-to-end deployment pipeline that transforms pretrained ViTs into a single runtime-configurable binary, enabling dynamic compute-budget switching on embedded CPUs. This is achieved by restructuring generated C kernels with modified loop bounds and binary-mask control logic, allowing execution to switch across discrete sparsity levels via compact external configuration files. Compared to multi-binary deployment, the proposed runtime-adaptive approach reduces on-device storage by up to 4.86x, requiring only 163 MB for ViT-Base instead of nearly 800 MB. To maximize pruning efficiency, we introduce a hardware-aligned block pruning strategy for Multi-Layer Perceptron (MLP) layers. In addition, a custom ISA extension is proposed to exploit input-reuse patterns in linear projection kernels. On a Synopsys TRV32P3FX RISC-V processor, the full system achieves up to 2.8x speedup at 65% MLP and 50% attention-head pruning for ViT-Base. The ISA extension alone provides a 1.56x speedup and 33% lower inference energy, with a 24.7% area overhead in a TSMC 28 nm implementation.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Vishnu PS, Ajay Kumar M, Yike Li, Robert Bogdan Staszewski, Deepu John. 2026-10-01. AdaptViT: Runtime-Adaptive Vision Transformer Deployment on Custom RISC-V. https://arxiv.org/abs/2610.02288

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Cross-Layer Analysis of Thermal Tuning Stalls in Wafer-Scale Optical Interconnects for LLM MoE Training

Mixture-of-experts (MoE) training depends heavily on all-to-all communication, which wafer-scale optical interconnects with dense wavelength-division multiplexing can serve. Their microring resonators rely on thermo-optic tuning, while MoE compute bursts swing the photonic-layer temperature by about 10 K within tens of milliseconds. This paper quantifies the communication stall by coupling packet-level network simulation, transient thermal simulation of a 3D-stacked GPU and photonic die, and a ring detuning criterion, and by feeding the stall back into the network timeline until both agree. Including this feedback, a tracking loop at the measured 5 nm/s lengthens the training iteration of Mixtral 8x7B and LLaMA-MoE 6.7B by factors of about 1.16 and 1.25 at full model depth. Measured H100 die temperatures match the modeled swing of 300 ms bursts within 15%. A loop slewing eight times faster, a heater driven at each kernel launch, or a dummy load near 75% of peak power removes the stall. An athermalized lithium-niobate ring with a non-volatile ferroelectric setpoint removes it with no holding power or fast loop, which makes it suitable for wafer-scale optical interconnects.

cs.AR↗

BigMoMo: Efficient Inference of Large-Scale MoE with Speculative Decoding on Mobile Devices

Mixture-of-Experts (MoE) models expand language model capacity on smartphones, but expert offloading remains constrained by limited DRAM capacity and costly data movement. Sequential token routing couples expert execution to fragmented flash reads and multistage NPU preparation, leaving sparse computation stalled on weight transfers. Each transfer serves few tokens before execution moves on. We exploit the multi-token verification window of speculative decoding to decouple expert movement from single-token execution, enabling weight reuse, contiguous flash reads, and load-compute overlap. We present \textsc{BigMoMo}, a mobile MoE runtime that exploits this window across the memory hierarchy. It prunes speculative branches and expert activations using acceptance rates, routing impact, and movement cost; reorganizes on-flash experts according to runtime co-loading patterns; and batches ready experts to overlap NPU computation with pending transfers. Across four MoE models and five benchmarks on two mobile platforms, \textsc{BigMoMo} achieves mean decoding speedups of $4.83\times$ over on-demand autoregressive offloading and $1.82\times$ over the best speculative MoE baseline, supporting MoE models up to 80B parameter.

cs.AR↗

Provisioning to Runtime Optimization of a 100 MW-Scale AI Cluster

The electric power supply for AI data centers is now the most significant bottleneck in the race toward Artificial General Intelligence, surpassing even the constraint of AI accelerator availability. To our knowledge, this paper is the first to describe the end-to-end power management process for a hyper-scale AI datacenter; from early power planning to accommodate next-generation accelerators 6--12 months before their general availability, to tuning power settings after large scale deployment, and finally to dynamic, runtime power management for evolving workloads. We present detailed power measurements for a 150 MW datacenter hosting a cluster of 83K GB200 GPUs. We share insights from building this state-of-the-art AI cluster. We hope this work encourages practitioners across the industry to share their own experiences as well.

cs.AR↗