Tokens per second (TPS) is a misleading metric for evaluating LLM performance. What actually matters is total task completion time — from start to finish. High TPS means nothing if the surrounding tooling is inefficient, the model can't find its footing, or the overall job takes longer than it should. A model spinning at 200 TPS but failing to complete tasks efficiently is just wasting cycles faster.
•1m watch time