17,000 Tokens/Second Inferencing of an AI Model 🤯
This title could be clearer and more informative.Try out Clickbait Shieldfor free (5 uses left this month).
Talis, a hardware startup, claims their custom silicon achieves ~17,000 tokens/second inference speed for the Llama 3.1 8B model — reportedly 10x faster than Nvidia H200/B200 GPUs at 20x lower infrastructure cost. A live demo shows ~15,500 tokens/second generating Playwright/Selenium C# test code nearly instantly. The video argues ultra-fast inference will dramatically accelerate AI agent workflows, potentially compressing 20-30 minute multi-agent tasks into seconds.
•6m watch time