Research7 min readSpeculative Decoding: The Inference Trick Hiding in Plain SightSpeculative decoding promises faster LLM inference without touching model weights. The math holds up — but the production story is messier than the papers admit.Aug 8, 2026Read