Search arXiv⌕ Search

arXiv · 2609.32055

Towards Simple Models of Complex SmartNICs

Abstract

Cloud vendors push ambitious in-network processing (e.g., crypto, telemetry) onto the NIC to offload servers even as link rates climb to terabit speeds. Vendors have responded with heterogeneous SmartNICs. For example, NVIDIA BlueField-3 interposes---between the wire and the host CPUs---a line-rate eSwitch, a multithreaded Data-Path Accelerator, general-purpose ARM cores, and a sea of fixed-function accelerators. These devices are notoriously hard to program, and harder still to predict. Applications can be implemented in many ways, with each choice potentially hitting a different bottleneck. A designer ideally needs to know---cheaply, and before a line of code is written---feasible choices and their bottlenecks, and design patterns to improve performance. Our paper offers a starting point to answer these questions using what we call the ZRAM model. It pairs a platform graph of processing zones and their channels with a program graph of tasks and their traffic fractions. The application designer or a compiler chooses a placement that maps the program graph onto the platform graph. Three metrics computed directly from this mapping---capability, roofline, and capacity---score the placement, deciding its feasibility and naming the bottleneck resource. We use a DDoS detector as a primary case study, and briefly explore two other applications, decision-tree inference and RDMA traversal. We distill seven design patterns for programming SmartNICs including a key one we call sifting. ZRAM generalizes to other SmartNICs such as Intel IPU E2200 and AMD Pensando Salina 400, and opens a new research agenda that includes compilers and hardware design.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Robert Chang, Teng Jiang, Wonsup Yoon, Daehyeok Kim, Sam Kumar, George Varghese. 2026-09-25. Towards Simple Models of Complex SmartNICs. https://arxiv.org/abs/2609.32055

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Space Filling Curves is All You Need: Communication-Avoiding Matrix Multiplication Made Simple

General Matrix Multiplication (GEMM) is the cornerstone of HPC workloads and Deep Learning. State-of-the-art (SOTA) vendor libraries tune tensor layouts, parallelization schemes and cache blocking to minimize data movement across the memory hierarchy and maximize throughput. However, optimal settings for these parameters depend on the target platform and matrix shapes, making exhaustive tuning infeasible. In this work, we address this cumbersome scheduling search and tuning using space-filling curves (SFC). We partition the matrix multiplication using advancements in SFC, and obtain platform-oblivious and shape-oblivious matrix multiplication schemes with a high degree of data locality. We extend the SFC-based work partitioning to implement Communication-Avoiding (CA) algorithms with replication techniques in a seamless fashion. The resulting SFC-CA GEMM achieves provable asymptotic communication optimality for both square and rectangular matrix regimes. Across four x86 and Arm platforms, SFC-CA GEMM outperforms vendor libraries by up to 5.5$\times$ per shape and 1.8$\times$ in weighted harmonic mean (WHM) throughput. Last, we show the impact of our work on two real-world applications by leveraging our SFC-CA GEMM as a compute backend: i) prefill of LLM inference with speedups up to 1.85$\times$ over SOTA inference runtimes, and ii) distributed-memory matrix multiplication with speedups up to 2.3$\times$ over the SOTA distributed-memory GEMM framework with vendor-optimized compute backend.

cs.DC↗

MoEless: Efficient MoE LLM Serving with Serverless Experts

Large Language Models (LLMs) increasingly adopt Mixture-of-Experts (MoE) architectures to scale efficiently under stringent resource constraints. However, MoE's sparse activation causes severe expert load imbalance, where a few experts become stragglers while others remain underutilized, leading to inflated inference latency and cost. Existing solutions assume static, serverful model deployments, limiting expert elasticity and often incurring costly expert swapping or degraded output quality. We present MoEless, an efficient serverless MoE serving framework that mitigates expert load imbalance via elastic expert execution. MoEless leverages lightweight, layer-aware predictors to estimate incoming expert load distributions and proactively identify stragglers. We design optimized scaling and placement strategies to improve function locality, GPU utilization, and cross-expert load balance. MoEless is prototyped on top of Megatron-LM and deployed on an eight-GPU testbed. Experiments with open-source MoE models and real-world workloads show that MoEless reduces inference latency by 43% and inference cost by 84% compared to state-of-the-art solutions.

cs.DC↗

Cascadia: A Control-Plane-Free Alternative to Hyperconverged AI Infrastructure

We present Cascadia, a system for serving large language models on fleets of commodity Intel AIPCs using their CPU, integrated-GPU, and NPU resources. Every node embeds ingress, scheduling, and execution; inference requests require no dedicated routing control plane. Nodes join a libp2p QUIC mesh using CA-issued ed25519 admission certificates, gossip signed capabilities, exchange live load over direct peer streams, and route OpenAI-compatible requests to eligible peers. An operator-run certificate authority handles admission and fleet management outside the inference path. Three serving modes share one interface: whole-model execution on one node, load-balanced replicas, and pipeline-sharded chains using the compilation and speculative decoding mechanism of our companion paper. Optional KV-cache mobility reuses compatible conversation prefixes after a routing move, with cold recomputation on a miss. Signed response receipts and hash-chained logs support provenance and audit. A three-node Phi-3.5-mini NPU testbed delivered 3.10x the response throughput of its one-node configuration under ten concurrent requests; a separate four-node deployment recorded 4.06x the throughput of direct single-node serving. Paired latency observations, runtime measurements, and internal functional checks characterize the tested configurations. We compare Cascadia with IBM, Nutanix, VMware, and HPE platforms on deployment footprint, hardware requirements, scheduling, scaling, licensing, and trust, using vendor documentation. The paper repository provides benchmark scripts, curated measurements, and a claim-to-evidence map.

cs.DC↗