This post introduces speculative decoding as an optimization technique for inference, specifically for large language models. It discusses the benefits of speculative decoding and how it can coexist with other optimization techniques. The post also provides code and resources for training speculators and showcases the speedup achieved in an internal production environment.
Table of contents
Speculative decoding: InferenceResultsSpeculative decoding: TrainingConclusion and Future Work27 Impressions