arXiv · 2610.05089
Cooldown Landmines: Cross-Tenant Interference Attacks on LLM Gateways
Abstract
LLM gateways enforce separate tenant quotas while sharing model deployments and cooldown records that temporarily exclude failing backends. However, a tenant's request failure can update these shared records and restrict other tenants' access to serviceable deployments. We identify two attacks that exploit this gap in LiteLLM. The first uses requests rejected at the key's requests-per-minute (RPM) limit: caller-supplied identifiers for known registered deployments reach failure handling, allowing two rejected requests to redirect another tenant to fallback with zero upstream calls from those requests. The second uses admitted traffic to create cooldown records that persist after backend capacity recovers. To address these failures, we design an origin check that blocks deployment updates from key RPM rejections and tenant-scoped cooldown that preserves the triggering tenant's back-off while retaining shared records for backend faults. Experiments with authenticated proxies and self-hosted vLLM demonstrate that the attacks can force fallback or denial while deployments remain serviceable. Across five paired four-worker runs, tenant scoping reduces victim fallback after recovery from 54/60 to 0/60, while increasing attempts against exhausted shared quotas. These findings show that tenant isolation must cover both request admission and the failure handling that governs shared deployment availability.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yudong Gao, Linghan Chen, Wenhan Wu, Quan Shi, Xutao Mao, Mia Zhou, Junjian Li, Xiaolong Liu, Jiyao Wang, Mingyu Guo, Honglong Chen. 2026-10-04. Cooldown Landmines: Cross-Tenant Interference Attacks on LLM Gateways. https://arxiv.org/abs/2610.05089
Cite the original work for its findings. Save a collection to share your selection of sources.