
Research7 min read
Speculative Decoding: The Inference Trick Hiding in Plain Sight
Speculative decoding promises faster LLM inference without touching model weights. The math holds up — but the production story is messier than the papers admit.
103 views
Read