One Threshold Is Not Enough: Prompt-Invariant Caching Schedules for Video Diffusion
Video diffusion models remain expensive because denoising repeatedly evaluates a costly model backbone. Dynamic caching methods like TeaCache, EasyCache, and DiCache reduce this cost by reusing prior outputs while an accumulated drift signal remains below a threshold $τ$. Although existing methods improve the drift signal, they hold $τ$ fixed throughout denoising. We show that fixed thresholds are structurally suboptimal: varying $τ$ across denoising achieves quality-latency tradeoffs unattainable by any fixed threshold across three caching methods with distinct drift signals. We further show that, surprisingly, each method's drift trajectory is nearly prompt-invariant across seven method-model combinations. The shape of the drift trajectory is determined by the model and caching method; therefore, a threshold schedule can be calibrated offline and reused across generations. Based on these findings, we present ACID (Adaptive Caching for vIDeo generation), which uses a low threshold during critical regions where drift changes rapidly and a high threshold across stable regions. ACID identifies these regions offline from the drift signal's second derivative, requires no training, and adds negligible runtime overhead. Across TeaCache, EasyCache, and DiCache on HunyuanVideo, Wan 2.1, and CogVideoX, ACID pushes the speed-quality Pareto frontier beyond fixed thresholds. On TeaCache with HunyuanVideo, it achieves 2.16x speedup over no caching and 38% additional speedup over a conservative fixed threshold, with less than 0.3 dB PSNR, 0.01 SSIM, and 0.01 LPIPS degradation relative to that fixed-threshold configuration.