Search arXiv⌕ Search

arXiv · 2609.38706

Preserving Provenance in Shared KV Caches for LLM Serving

Abstract

Production LLM serving stacks combine an inference engine's local prefix cache with a shared KV-cache tier for fleet-wide reuse. The local cache distinguishes requests by adapter, weight configuration and sharing domain, but the shared tier may key entries only by token content and coarse model metadata. This boundary erases provenance and lets identical tokens under incompatible computational or sharing contexts collide. We call this composition gap provenance-blind reuse and present its first systematic study. A source audit of three vLLM connectors confirms the structural omission, while runtime experiments reproduce it across vLLM and two SGLang releases, 12 models from 7 families (0.5 B-32 B), and over 160 configurations. Cross-adapter collisions reduce accuracy from 0.94 to 0.64, incompatible KV representations reduce reasoning accuracy to zero, and salt omission enables 93% prompt identification from timing. We formalize the missing guarantee as the KV provenance contract: for a declared dimension registry, shared keys must be injective over computational and sharing provenance, with identities stable across workers. Any dimension with a stable identity can therefore be added without connector-specific key logic. A canonical descriptor binds per-request and per-worker provenance into lookup and store keys, while a differential checker detects dimensions that change KV state without changing the key. Implemented in vLLM and SGLang 0.5.20 across three cache paths, provenance binding eliminates unsafe reuse while preserving legitimate sharing. Hit-path latency changes remain within 0.34 ms and below run-to-run variation; retention grows with provenance diversity.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Wei Song, Yuxin Cao, Xi Zheng, Leo Zhang, Xiao Cheng. 2026-09-30. Preserving Provenance in Shared KV Caches for LLM Serving. https://arxiv.org/abs/2609.38706

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Space Filling Curves is All You Need: Communication-Avoiding Matrix Multiplication Made Simple

General Matrix Multiplication (GEMM) is the cornerstone of HPC workloads and Deep Learning. State-of-the-art (SOTA) vendor libraries tune tensor layouts, parallelization schemes and cache blocking to minimize data movement across the memory hierarchy and maximize throughput. However, optimal settings for these parameters depend on the target platform and matrix shapes, making exhaustive tuning infeasible. In this work, we address this cumbersome scheduling search and tuning using space-filling curves (SFC). We partition the matrix multiplication using advancements in SFC, and obtain platform-oblivious and shape-oblivious matrix multiplication schemes with a high degree of data locality. We extend the SFC-based work partitioning to implement Communication-Avoiding (CA) algorithms with replication techniques in a seamless fashion. The resulting SFC-CA GEMM achieves provable asymptotic communication optimality for both square and rectangular matrix regimes. Across four x86 and Arm platforms, SFC-CA GEMM outperforms vendor libraries by up to 5.5$\times$ per shape and 1.8$\times$ in weighted harmonic mean (WHM) throughput. Last, we show the impact of our work on two real-world applications by leveraging our SFC-CA GEMM as a compute backend: i) prefill of LLM inference with speedups up to 1.85$\times$ over SOTA inference runtimes, and ii) distributed-memory matrix multiplication with speedups up to 2.3$\times$ over the SOTA distributed-memory GEMM framework with vendor-optimized compute backend.

cs.DC↗

MoEless: Efficient MoE LLM Serving with Serverless Experts

Large Language Models (LLMs) increasingly adopt Mixture-of-Experts (MoE) architectures to scale efficiently under stringent resource constraints. However, MoE's sparse activation causes severe expert load imbalance, where a few experts become stragglers while others remain underutilized, leading to inflated inference latency and cost. Existing solutions assume static, serverful model deployments, limiting expert elasticity and often incurring costly expert swapping or degraded output quality. We present MoEless, an efficient serverless MoE serving framework that mitigates expert load imbalance via elastic expert execution. MoEless leverages lightweight, layer-aware predictors to estimate incoming expert load distributions and proactively identify stragglers. We design optimized scaling and placement strategies to improve function locality, GPU utilization, and cross-expert load balance. MoEless is prototyped on top of Megatron-LM and deployed on an eight-GPU testbed. Experiments with open-source MoE models and real-world workloads show that MoEless reduces inference latency by 43% and inference cost by 84% compared to state-of-the-art solutions.

cs.DC↗

Cascadia: A Control-Plane-Free Alternative to Hyperconverged AI Infrastructure

We present Cascadia, a system for serving large language models on fleets of commodity Intel AIPCs using their CPU, integrated-GPU, and NPU resources. Every node embeds ingress, scheduling, and execution; inference requests require no dedicated routing control plane. Nodes join a libp2p QUIC mesh using CA-issued ed25519 admission certificates, gossip signed capabilities, exchange live load over direct peer streams, and route OpenAI-compatible requests to eligible peers. An operator-run certificate authority handles admission and fleet management outside the inference path. Three serving modes share one interface: whole-model execution on one node, load-balanced replicas, and pipeline-sharded chains using the compilation and speculative decoding mechanism of our companion paper. Optional KV-cache mobility reuses compatible conversation prefixes after a routing move, with cold recomputation on a miss. Signed response receipts and hash-chained logs support provenance and audit. A three-node Phi-3.5-mini NPU testbed delivered 3.10x the response throughput of its one-node configuration under ten concurrent requests; a separate four-node deployment recorded 4.06x the throughput of direct single-node serving. Paired latency observations, runtime measurements, and internal functional checks characterize the tested configurations. We compare Cascadia with IBM, Nutanix, VMware, and HPE platforms on deployment footprint, hardware requirements, scheduling, scaling, licensing, and trust, using vendor documentation. The paper repository provides benchmark scripts, curated measurements, and a claim-to-evidence map.

cs.DC↗