arXiv · 2609.40098
Pricing Time, Not Just Tokens: Latency-Aware Mechanism Design for LLM Inference
Abstract
The economic theory of LLM pricing treats tokens as a homogeneous commodity considering aggregate token count as the main features buyers and sellers consider. We model inference as a service market where buyers have three-dimensional private information - willingness-to-pay, task volume, and time preference - and utility depends on latency slack alongside token quantities. Our main result is a separation theorem: discrete hardware tiers induce endogenous self-selection on time preferences, reducing three-dimensional screening to standard one-dimensional screening within each tier. We derive the cost structure from GPU inference physics - compute-bound prefill and bandwidth-bound decode - and characterize optimal tiered mechanisms via virtual-value techniques. Optimal per-task prices are volume-independent, providing theoretical grounding for flat per-token API pricing. We verify the mechanism empirically by calibrating to 8-GPU clusters of H100 and B200 hardware. The separation theorem holds in 83% of 105 tested configurations overall, rising to 96% at economically relevant WTP scales. A seller adopting two-tier pricing under the optimal mechanism captures 26-66% higher profit than the best single-tier alternative, with gains driven by efficient cross-tier allocation in regimes where hardware costs are a significant fraction of per-request value.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Ian McDougall, Karthikeyan Sankaralingam. 2026-09-30. Pricing Time, Not Just Tokens: Latency-Aware Mechanism Design for LLM Inference. https://arxiv.org/abs/2609.40098
Cite the original work for its findings. Save a collection to share your selection of sources.