arXiv · 2610.04951
LLBPE: Linked-List Based GPU-Parallel BPE Tokenizer
Abstract
Every LLM inference begins with tokenization, which converts raw input bytes into the discrete token sequence the model consumes. For text, this step is often implemented using Byte Pair Encoding (BPE), an algorithm originally introduced for data compression. BPE has traditionally run on the CPU with extensive optimization, but recent work has moved it to the GPU for higher throughput. We show that these GPU implementations are bottlenecked not by computation but by data movement. We develop LLBPE that represents the token sequence as an array-based linked list so that each merge reduces to a constant- time pointer update. Furthermore, LLBPE fuses rank lookup, minimum selection, and merging into a single kernel to eliminate redundant hash map queries. LLBPE achieves up to 5.2x higher throughput than the best existing GPU implementation and 24.6x over optimized CPU implementations, at the cost of minor discrepancies in tokenized output.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Aditya Kovilur, Varun Chandra Shekar, Ariful Azad. 2026-10-04. LLBPE: Linked-List Based GPU-Parallel BPE Tokenizer. https://arxiv.org/abs/2610.04951
Cite the original work for its findings. Save a collection to share your selection of sources.