Search arXiv⌕ Search

arXiv · 2609.40098

Pricing Time, Not Just Tokens: Latency-Aware Mechanism Design for LLM Inference

Abstract

The economic theory of LLM pricing treats tokens as a homogeneous commodity considering aggregate token count as the main features buyers and sellers consider. We model inference as a service market where buyers have three-dimensional private information - willingness-to-pay, task volume, and time preference - and utility depends on latency slack alongside token quantities. Our main result is a separation theorem: discrete hardware tiers induce endogenous self-selection on time preferences, reducing three-dimensional screening to standard one-dimensional screening within each tier. We derive the cost structure from GPU inference physics - compute-bound prefill and bandwidth-bound decode - and characterize optimal tiered mechanisms via virtual-value techniques. Optimal per-task prices are volume-independent, providing theoretical grounding for flat per-token API pricing. We verify the mechanism empirically by calibrating to 8-GPU clusters of H100 and B200 hardware. The separation theorem holds in 83% of 105 tested configurations overall, rising to 96% at economically relevant WTP scales. A seller adopting two-tier pricing under the optimal mechanism captures 26-66% higher profit than the best single-tier alternative, with gains driven by efficient cross-tier allocation in regimes where hardware costs are a significant fraction of per-request value.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Ian McDougall, Karthikeyan Sankaralingam. 2026-09-30. Pricing Time, Not Just Tokens: Latency-Aware Mechanism Design for LLM Inference. https://arxiv.org/abs/2609.40098

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Throttling Equilibria in Auction Markets

Throttling is a popular method of budget management for online ad auctions in which the platform modulates the participation probability of an advertiser in order to smoothly spend her budget across many auctions. In this work, we investigate the setting in which all of the advertisers simultaneously employ throttling to manage their budgets, and we do so for both first-price and second-price auctions. We analyze the structural and computational properties of the resulting equilibria. For first-price auctions, we show that a unique equilibrium always exists, is well-behaved and can be computed efficiently via tatonnement-style decentralized dynamics. In contrast, for second-price auctions, we prove that even though an equilibrium always exists, the problem of finding even an approximate equilibrium is PPAD-complete, there can be multiple equilibria, and it is NP-hard to find the revenue maximizing one. We also compare the equilibrium outcomes of throttling to those of multiplicative pacing, which is the other most popular and well-studied method of budget management. Finally, we characterize the Price of Anarchy of these equilibria for liquid welfare by showing that it is at most 2 for both first-price and second-price auctions, and demonstrating that our bound is tight.

cs.GT↗

Strategic Disclosure of Action Space in Principal-Agent Contracts

We study strategic disclosure of the action space in principal-agent contracting, where an agent selects a disclosed action set to shape the principal's perception of her capabilities before contract design. Unaware of the strategic disclosure, the principal designs a revenue-optimal contract as if the disclosed action set were complete and accurate. We consider two variants distinguished by cost verifiability. When costs are unverifiable, the agent can extract the entire first-best surplus, leaving the principal with zero revenue. When costs are verifiable, we characterize the agent's optimal disclosure strategy in binary-outcome settings and, more generally, when the principal is restricted to linear contracts, reducing the agent's problem to a two-variable convex optimization problem. We prove that the agent can secure utility of at least a $1/e$ fraction of the first-best surplus, which also yields a $1/e$ welfare guarantee under optimal disclosure. While the principal's revenue can be arbitrarily small compared to the first-best surplus, when the ratio of maximum to minimum expected reward among non-null base actions is at most $L$, we establish a revenue guarantee of $Θ(1/\log L)$ relative to the first-best surplus. We also compare utilities and welfare under strategic disclosure with their counterparts in the canonical model. Finally, we extend the agent's $1/e$ utility guarantee to general outcome spaces without restricting the principal to linear contracts. Our results show how strategic action-space disclosure changes the distribution of surplus while preserving a constant-factor welfare guarantee under the agent's optimal disclosure.

cs.GT↗

On the Geographic Incentives of Multiple Concurrent Proposers

We study whether systems with multiple concurrent proposers (MCP) incentivize the geographic decentralization of proposers. In our model, concurrent proposers strategically choose their locations to maximize their captured market share of users, who decide where to send transactions by comparing each proposer's time to inclusion. The time to inclusion between a user and a proposer consists of two terms: the distance between them, and the distance between the proposer and her nearest quorum of attesters, the latter of which we term the quorum radius. We formally prove that proposers' quorum radii are the determinant of their incentives: proposers optimally choose where to locate in order to minimize their quorum radii. We obtain as a result co-location when proposers also serve as attesters, but geographic diversity when these roles are separate. We also examine pre-confirmations and faster intervalidator connections, showing that they attenuate the influence of the quorum radius and further destabilize co-location incentives.

cs.GT↗