TriForce is a hierarchical speculative decoding system designed for efficient serving of large language models (LLMs) in long contexts. It addresses the bottlenecks of KV cache and model weights, resulting in significant speedups.
1 Impression
TriForce is a hierarchical speculative decoding system designed for efficient serving of large language models (LLMs) in long contexts. It addresses the bottlenecks of KV cache and model weights, resulting in significant speedups.