Search arXiv⌕ Search

arXiv · 2610.03045

PEEK: Heterogeneous Parallelism for Privileged Error Detection in Safety-Critical Processors

Abstract

Heterogeneous parallel error detection architecture has been widely studied for safeguarding OoO superscalar processors in safetycritical systems, as it achieves significantly lower hardware overhead compared to traditional LockStep, by exploiting the parallelism that exists in a secondary execution. However, previous works do not cover the protection of privileged-mode execution, impeding their effectiveness in real-world deployment. Moreover, naive extension to privileged-mode can cause a litany of issues, from abysmal performance due to high synchronization costs, to full deadlocks. Here, we present PEEK, the first privileged parallel error detection architecture. Based on a deep analysis of privileged execution, we redesign the verification pipeline, addressing all the bottlenecks and bugs identified in privileged-mode protection. Evaluated using various metrics on an RTL-level full system running Linux, PEEK achieves full-privilege protection on Linux with negligible performance slowdown and affordable hardware overhead. PEEK has been taped out using a 28nm process, and its source is available at https://anonymous.4open.science/r/PEEK-3000.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Tinglue Wang, Zhenghui Guo, Bing Guo, Jiapeng Guan, Renshuang Jiang, Jie Zhang, Jing Li, Xin Si, Sam Ainsworth, Zhe Jiang. 2026-10-02. PEEK: Heterogeneous Parallelism for Privileged Error Detection in Safety-Critical Processors. https://arxiv.org/abs/2610.03045

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Cross-Layer Analysis of Thermal Tuning Stalls in Wafer-Scale Optical Interconnects for LLM MoE Training

Mixture-of-experts (MoE) training depends heavily on all-to-all communication, which wafer-scale optical interconnects with dense wavelength-division multiplexing can serve. Their microring resonators rely on thermo-optic tuning, while MoE compute bursts swing the photonic-layer temperature by about 10 K within tens of milliseconds. This paper quantifies the communication stall by coupling packet-level network simulation, transient thermal simulation of a 3D-stacked GPU and photonic die, and a ring detuning criterion, and by feeding the stall back into the network timeline until both agree. Including this feedback, a tracking loop at the measured 5 nm/s lengthens the training iteration of Mixtral 8x7B and LLaMA-MoE 6.7B by factors of about 1.16 and 1.25 at full model depth. Measured H100 die temperatures match the modeled swing of 300 ms bursts within 15%. A loop slewing eight times faster, a heater driven at each kernel launch, or a dummy load near 75% of peak power removes the stall. An athermalized lithium-niobate ring with a non-volatile ferroelectric setpoint removes it with no holding power or fast loop, which makes it suitable for wafer-scale optical interconnects.

cs.AR↗

BigMoMo: Efficient Inference of Large-Scale MoE with Speculative Decoding on Mobile Devices

Mixture-of-Experts (MoE) models expand language model capacity on smartphones, but expert offloading remains constrained by limited DRAM capacity and costly data movement. Sequential token routing couples expert execution to fragmented flash reads and multistage NPU preparation, leaving sparse computation stalled on weight transfers. Each transfer serves few tokens before execution moves on. We exploit the multi-token verification window of speculative decoding to decouple expert movement from single-token execution, enabling weight reuse, contiguous flash reads, and load-compute overlap. We present \textsc{BigMoMo}, a mobile MoE runtime that exploits this window across the memory hierarchy. It prunes speculative branches and expert activations using acceptance rates, routing impact, and movement cost; reorganizes on-flash experts according to runtime co-loading patterns; and batches ready experts to overlap NPU computation with pending transfers. Across four MoE models and five benchmarks on two mobile platforms, \textsc{BigMoMo} achieves mean decoding speedups of $4.83\times$ over on-demand autoregressive offloading and $1.82\times$ over the best speculative MoE baseline, supporting MoE models up to 80B parameter.

cs.AR↗

Provisioning to Runtime Optimization of a 100 MW-Scale AI Cluster

The electric power supply for AI data centers is now the most significant bottleneck in the race toward Artificial General Intelligence, surpassing even the constraint of AI accelerator availability. To our knowledge, this paper is the first to describe the end-to-end power management process for a hyper-scale AI datacenter; from early power planning to accommodate next-generation accelerators 6--12 months before their general availability, to tuning power settings after large scale deployment, and finally to dynamic, runtime power management for evolving workloads. We present detailed power measurements for a 150 MW datacenter hosting a cluster of 83K GB200 GPUs. We share insights from building this state-of-the-art AI cluster. We hope this work encourages practitioners across the industry to share their own experiences as well.

cs.AR↗