Search arXiv⌕ Search

arXiv · 2609.34663

Tool Waiting and Re-arrival in Compile-Time-Static LLM Serving: Cost Mechanisms and Configuration Selection

Abstract

In agentic LLM services, a session calls an external tool, waits for it, and re-arrives to continue inference. Statically compiled NPU serving can fix the batch bucket set, the maximum batch size, and the number of KV cache slots at compile time. We define such an environment as a compile-time-static serving substrate and analyze the execution-time cost that tool waiting and re-arrival incur in it. On a single LLM instance, we run synthetic workloads following a measured tool waiting time distribution and compare, on the same inputs, a baseline configuration with settings {1, 2, 4, 8}, 8, and 8 against configurations that change some of them. Because re-arrival times differ across configurations, we build a simulator that replays request processing in time order, select the candidate with the lowest predicted cost among 2,077 configurations, and validate it on new inputs. We identify three mechanisms: discrete batch alignment, KV cache survival, and prefill interference. At a concurrency of 6, absent from the bucket set, tool waiting lowered the padding ratio (0.235 to 0.120) yet increased decode execution time 1.51-fold, so padding alone did not indicate cost. On new inputs at a concurrency of 8, where the baseline reused KV in 9 of 24 re-arrivals, enlarging the maximum batch size alone cut execution cost by 8.25%, and the selected configuration, which also adjusted the bucket set, by 9.72%. Where 17 of 18 re-arrivals were already reused, the effect was 0.59%. Compile-time configurations should thus be selected by diagnosing KV reuse loss and the resulting change in execution.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Dongkyeom Jang, In-Nea Wang, Junho Jeong. 2026-09-28. Tool Waiting and Re-arrival in Compile-Time-Static LLM Serving: Cost Mechanisms and Configuration Selection. https://arxiv.org/abs/2609.34663

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Accelerating Transfer-Learning-Based Autotuning with Predictive LLVM IR Performance Ranking

As the complexity of High Performance Computing (HPC) ecosys- tems continually increases, achieving optimal performance becomes a challenge. Traditional performance autotuning techniques pro- vide promising means to navigate this complexity, these techniques remain computationally intensive and require many evaluations to find optimal configurations. This work proposes an autotuning framework that designs a machine learning-based ensemble LLVM Intermediate Representa- tion (IR) ranker, Neural Configuration Scorer (NCS). NCS ranks the performance of IRs sampled by a transfer-learning-based autotuner, improving the efficiency of the tuning process by reducing tuning overheads and circumventing subpar evaluations. By leveraging knowledge from related tasks, we are able to effectively exploit the transfer relationship to access high-performing configurations in fewer samples than traditional techniques that rely upon itera- tive refinement. Our framework can achieve similar performance improvements as state-of-the-art autotuning techniques with up to 61.67% fewer evaluations, averaging 27.85% fewer evaluations across various HPC benchmarks.

cs.PF↗

PASCAL: A Progress Divergence-Aware Shared-Cache Model

In modern AI accelerators and GPUs, many concurrent cores repeatedly access the same shared data. This pattern occurs in attention, where different query (Q) tiles share the same key and value (K/V) blocks, GEMM, where every tile in a row reads the same slice, and many other operators. We name this pattern shared cyclic scan. Due to a significant amount of data reuse in this pattern, the cache is expected to capture as much data reuse as possible and largely reduce requests sent to the main memory for both performance and energy consumption concerns. However, in reality, because of the intrinsic asynchrony of multi-cores, the actual cache miss rate and DRAM traffic can be much higher compared to ideal cases. In this paper, we propose PASCAL, a shared-cache model for shared cyclic scans. It calibrates finite-run traffic, which reflects progress divergence, on reference configurations and interpolates it along static program structure such as occupancy to predict the miss rate before execution. PASCAL supports software configuration exploration without target traces or counters at scales where cycle-accurate simulation is impractical, while its analysis gives a sharp $2σ-1$ sufficient capacity condition for preserving LRU sharing. Across held-out scans and GEMM on GB10 and Thor, PASCAL reaches 12.04% balanced fill-equivalent miss-rate MAPE. Replacing TileSight's cache component with PASCAL lowers GB10 GEMM latency MAPE from 18.17% to 13.14%.

cs.PF↗

Bandwidth, Not FLOPS: FFT Kernels, Matrix Units and SAR Imaging on Apple M6

The fast Fourier transform (FFT) underlies radar, imaging and scientific computing. A classic rule for fast GPU FFTs is to compute the largest block that fits on chip and compose larger transforms from such blocks. We test this rule on Apple's M6 chip, whose GPU performs 29 arithmetic operations in the time it reads one byte from main memory, and which adds matrix units to both its GPU and CPU. Comparing our kernels with Apple's MPSGraph and vDSP libraries and MLX, we find that data movement, not arithmetic, sets FFT speed. Large batches run at or near the main-memory bandwidth limit in every GPU library, and benchmarks that keep data in cache overstate real throughput by up to $3.7\times$. The on-chip rule still predicts where speed collapses: MPSGraph and MLX lose half or more of their speed once a transform outgrows the GPU's 32\,KiB local memory. Keeping such transforms in registers avoids an extra trip through memory: $2.2\times$ faster than MPSGraph, and $4.4\times$ with half-precision storage. The GPU's matrix units do not help, because recasting the FFT as matrix products adds as much arithmetic as they save. The CPU's matrix unit, 47 times faster than its vector units, does: our kernel for it beats vDSP by up to $5.3\times$. End to end, a $4096\times4096$ synthetic aperture radar image takes 8.1\,ms on the GPU (6.0\,ms with half-precision intermediates), $14$--$19\times$ faster than a CPU reference. Running the CPU's matrix unit alongside the GPU gains nothing: both share one memory bandwidth.

cs.PF↗