Search arXivSearch

arXiv · 2608.00897

Shiftfly: Scaling the Accelerator Interconnect Past the Pod with a Shift-Routed Optical Tier

Abstract

Google's TPU interconnect spent nine generations as a $k$-ary $n$-cube, whose diameter grows as $Θ(N^{1/n})$, before TPU 8i replaced it with Boardfly: a three-tier hierarchy in which every tier is a complete graph, giving pod diameter 7 over 1,152 chips. Boardfly suits a single inference pod but does not extend, because a complete global tier needs $G-1$ optical ports to reach $G$ groups. A 400,000-chip machine would need 12,499 per group against the 40 available, so another hierarchy level must be stacked, and each level costs four chip hops. We propose Shiftfly, which keeps Boardfly's building block and group verbatim and replaces only the global tier with a generalized Kautz digraph, installed as a fixed permutation on the optical circuit switch the fabric already owns. At an identical 40 ports per group, Shiftfly is flat, has guaranteed diameter $\lceil \log_d G \rceil$, routes without tables by a shift register, and addresses content natively. The evaluation is deliberately two-sided. Shiftfly loses at one-pod scale, where Boardfly achieves chip-level diameter 7 against Shiftfly's 11, and wins beyond it, cutting worst-case distance from 23 to 19 hops at 400,000 chips with roughly 2.7x better spectral expansion at equal cost. The redundancy argument that motivated the design does not survive its own evaluation: locality-aware placement supplies most of the achievable saving on shared content, leaving the shift algebra a 2.9% residual, and we report the metric inversion that conceals this. We also measure operability. Replacing a failed group costs the same 40 optical circuits in both designs; deriving the global wiring costs 28 bits of control-plane state against 550,000. Slice allocation is the one regression: an arbitrary induced subset of a shift graph is disconnected, so slices must be instantiated rather than carved.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Eylon E. Krause. 2026-08-01. Shiftfly: Scaling the Accelerator Interconnect Past the Pod with a Shift-Routed Optical Tier. https://arxiv.org/abs/2608.00897

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

CSI Simulation: Why Additive Noise Fails and How to Fix It

Channel State Information (CSI) has become a widely used wireless channel sensing modality for applications such as indoor localization, activity recognition, and respiration monitoring. Because collecting labeled data under every target condition is impractical, training CSI-based models often relies on simulated data produced by adding noise or perturbations to recorded channel estimates, most commonly additive white Gaussian noise (AWGN). This practice assumes that the receiver chain between the antenna and the channel estimator is linear and gain-invariant. We test this assumption empirically using RF jamming as a controlled perturbation on 6 commodity receivers across 2 indoor environments. The assumption does not hold. Automatic gain control compresses the channel estimate multiplicatively before digitization, producing amplitude distributions that no additive noise variance can reproduce. To close the resulting fidelity gap, we propose M_QTC, a measurement-calibrated model that learns the per-subcarrier distribution transformation through quantile mapping, temporal filtering, and copula-based cross-subcarrier reordering. M_QTC reduces amplitude error 8-fold and closes 89% of the aggregate fidelity gap across four complementary dimensions. The improvement transfers directly to downstream tasks, where 5 classifiers from different families trained on M_QTC-simulated data recover 93% of real-data jamming detection performance, while AWGN-trained classifiers remain near random decision.

cs.NI

BALANCE: Hybrid Autoregressive-Speculative LLM Inference at the Network Edge

Edge inference is a promising paradigm to provide large language model (LLM) inference services in next-generation mobile networks. LLM inference mainly relies on two approaches: Autoregressive decoding (AD) generates output tokens sequentially, resulting in long latency; Speculative decoding (SD) accelerates inference by using a small language model (SLM) to generate multiple draft tokens for LLM verification, but incurs extra memory costs. Due to this latency-memory tradeoff, neither approach alone can efficiently serve users with heterogeneous demands under limited edge computing resources. To address this challenge, we propose a hybrid autoregressive-speculative inference (BALANCE) framework for edge LLM inference. In BALANCE, an edge server hosts both an SLM and an LLM, admits users, assigns each admitted user to the AD or SD mode, and performs the two modes simultaneously. To maximize the number of served users, we formulate a task throughput maximization problem to jointly determine user admission and computing resource allocation between AD and SD under user latency requirements and server memory constraints. Since the problem is NP-hard, we develop a polynomial-time algorithm that transforms the original problem into two sub-problems and obtains a sub-optimal solution with a constant approximation guarantee. Experiments demonstrate that BALANCE consistently outperforms conventional AD and SD and significantly improves task throughput.

cs.NI

Zero-Knowledge Remote Adversarial Attack against Wi-Fi-based Human Activity Recognition for Privacy Protection

The growing capability of Wi-Fi devices to identify human activities using channel state information (CSI) raises privacy concerns. To counter this threat, we propose GRAW, an adversary system, acting as a privacy defender, that degrades the human activity recognition (HAR) system at the user device by perturbing the router's signals that the device uses to estimate CSI. GRAW employs generative adversarial imitation learning (GAIL) to construct perturbation signals, and thereby eliminates the need for any information on the target HAR systems and their inputs (i.e., zero-knowledge operation). We evaluate GRAW against seven representative HAR models, using datasets collected in five environments, including our own dataset. We observe that GRAW is the only remote attack scheme that degrades every tested HAR model to a random-selection level. At the same perturbation level, GRAW achieves an attack success ratio up to 76.7% higher than comparison methods, while maintaining over 99% packet success rate on regular Wi-Fi communication. We demonstrate the feasibility of GRAW through real-time, over-the-air experiments with software-defined radios.

cs.NI