Skip to Main Content

Draft-then-verify speculative decoding achieves 2-3x faster token generation with identical output quality, reducing GPU.Leviathan et al., 'Fast Inference from Transformer…