DFlash Speculative Decoding Drafts Whole Token Blocks in Parallel for Up to 15x Higher Throughput on NVIDIA Blackwell
Autoregressive giant language fashions generate textual content one token at a time. Each token waits for the one earlier than it. This serial loop leaves trendy GPUs underused and retains inference gradual. The value grows worse with lengthy Chain-of-Thought reasoning fashions. Their prolonged outputs make latency the dominant a part of era. Speculative decoding is…
