Search arXiv⌕ Search

arXiv · 2610.04951

LLBPE: Linked-List Based GPU-Parallel BPE Tokenizer

Abstract

Every LLM inference begins with tokenization, which converts raw input bytes into the discrete token sequence the model consumes. For text, this step is often implemented using Byte Pair Encoding (BPE), an algorithm originally introduced for data compression. BPE has traditionally run on the CPU with extensive optimization, but recent work has moved it to the GPU for higher throughput. We show that these GPU implementations are bottlenecked not by computation but by data movement. We develop LLBPE that represents the token sequence as an array-based linked list so that each merge reduces to a constant- time pointer update. Furthermore, LLBPE fuses rank lookup, minimum selection, and merging into a single kernel to eliminate redundant hash map queries. LLBPE achieves up to 5.2x higher throughput than the best existing GPU implementation and 24.6x over optimized CPU implementations, at the cost of minor discrepancies in tokenized output.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Aditya Kovilur, Varun Chandra Shekar, Ariful Azad. 2026-10-04. LLBPE: Linked-List Based GPU-Parallel BPE Tokenizer. https://arxiv.org/abs/2610.04951

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Energy-performance tradeoffs in server farms with batch services and setup times

Data centers consume a large amount of energy, much of which is wasted due to idle servers. Turning off idle servers might be an effective power-saving solution; however, there is a trade-off between energy savings and system performance. Hence, we propose a setup queueing model with a batching policy that allows servers to process a set of jobs simultaneously to minimize power consumption while maintaining acceptable performance. We consider an M/M/c/SET--BATCH queue, a multi-server batch service queue with a fixed batch size and setup times, and some variants, including systems in which idle servers delay before turning off or systems in which the batch size is dynamic. We analyze the steady-state probabilities and system performance of the M/M/c/SET--BATCH system and its variants. Our analysis of the M/M/c/SET--BATCH system with lower computational complexity is made possible by utilizing the special structure of the model. In addition, we use simulations to compare the M/M/c/SET--BATCH model with some other variants with different setup time distributions. The results suggest that the model performs better when the setup time has a larger coefficient of variation. Our results indicate that the batching policy enhances the system performance, especially when we allow servers to be idle before turning them off.

cs.PF↗

PonyTail: Profiling and Improving Service Tail Latency

Datacenter services must meet tight tail latency service-level objectives. Research on improving tail latency focuses on reducing queuing delay by approximating optimal request scheduling policies. As request scheduling nears optimality, the primary way to further improve tail latency is to reduce the service time of requests. Unfortunately, datacenter services lack obvious hotspots for general service time optimization. We point out a new optimization opportunity in targeting the tail service time of services with dispersive service times, which are common in datacenters. To show the feasibility and benefit of this approach, we present PonyTail, a performance analysis methodology and toolkit for analyzing service-level outlier behaviors associated with tail service time. PonyTail helps performance engineers identify control-flow and microarchitectural patterns responsible for service time outliers, as well as their root causes. We apply PonyTail to analyze five latency-critical services, including an in-memory database system and an online search engine. Based on PonyTail's analysis, we optimize some of the services, improving their 99th percentile service time and latency by 10.7%--46% and 21%--$5\times$, respectively.

cs.PF↗

ServeTwin: A Benchmark-Validated Simulator for Distributed LLM Architecture Exploration

Evaluating distributed LLM serving designs on physical clusters is costly. Yet existing simulators provide only subsets of the capabilities needed for realistic design exploration: stateful closed-loop execution, timing prediction without profiling target hardware, and direct execution of unmodified serving benchmarks. We present ServeTwin, a closed-loop simulator that couples specification-driven analytical timing with a stateful serving loop. This coupling captures feedback among request completion, scheduling, queue state, and KV-cache evolution, including prefill-decode disaggregation and multi-turn agent workloads. ServeTwin avoids target-hardware operator profiling through iSTAGE, an analytical trace generator that derives per-iteration traces from model and engine specifications while representing ragged batches and discrete engine mechanisms such as CUDA-graph batch padding. It further decomposes execution time into component-owned throughput, scheduler, and runtime costs. This ownership lets unaffected parameters transfer across platforms and confines recalibration to changed hardware or software components. ServeTwin implements a vLLM-compatible interface and runs unmodified serving benchmarks. Against real deployments, it reproduces InferenceX's steady-state throughput-interactivity frontier with a 3.6% mean error and predicts LMBenchmark's multi-turn performance with a 9.9% error while tracking KV-cache evolution. Pre-silicon sweeps reveal that the preferred HBM bandwidth-capacity tradeoff reverses across workload states and that increasing concurrency shifts the bottleneck from memory bandwidth to scheduler and runtime overheads. Together, these capabilities enable practical exploration of distributed LLM serving systems before target hardware is available.

cs.PF↗