Evaluating Losslessness in Speculative Decoding Under Finite-Precision Inference
Lossless speculative decoding is typically defined at the algorithmic level: a speculative procedure proposes multiple tokens and a verification procedure is designed to preserve the output trajectory of an autoregressive reference model exactly. In practical neural inference, however, this guarantee is implemented using finite-precision floating-point computations, and discrete token selection can amplify small numerical differences into divergent generation trajectories. We investigate this distinction using Orthrus, a hybrid autoregressive-diffusion architecture that performs self-drafting and self-verification within a frozen autoregressive backbone, as a representative case study. Across 1,190 prompts from 12 domains, exact trajectory matching under BF16 occurs for only 45\% of the authors' checkpoint generations and 43% of those from our independently trained model. The probability of matching is strongly associated with the response-conditional perplexity of the autoregressive reference, indicating that trajectory divergence is not uniform across inputs. Despite these divergences, Orthrus does not exhibit systematic degradation on the evaluated downstream tasks. In contrast, FP32 inference yields exact trajectory matching on all evaluated prompts. These results demonstrate a gap between algorithmic losslessness and its implementation under finite-precision arithmetic, and motivate evaluating lossless speculative decoding at the level of exact generation trajectories as well as downstream task performance.