Sequoia is a scalable, robust, and hardware-aware speculative decoding framework that enables serving LLMs on consumer GPUs with low latency. It can serve a Llama2-70B on a single RTX-4090 8 times faster than other offloading serving systems. Sequoia is scalable and robust, allowing for faster growth in accepted tokens and generating temperatures more effectively compared to alternative methods.
28 Impressions