Speculation Is All You Need!

This title could be clearer and more informative.Try out Clickbait Shieldfor free (5 uses left this month).

Speculative decoding can speed up LLM inference beyond the typical 2-3x cap. Modal's new DFlash draft models replace the standard autoregressive drafter with a block diffusion model that denoises a full token block in one parallel pass, and feeds the target model's hidden states into the drafter for better alignment. This pushed acceptance length from ~3 to over 9 tokens, achieving 4x+ speedups — over 1000 tokens/sec on Qwen 3.5 122B-A10B on a B200 GPU versus 250 without speculation. The post also covers ART (Agent Reinforcement Trainer), an open-source framework using GRPO to train LLM agents on multi-step tasks without manual reward engineering, with integrations for LangGraph, CrewAI, vLLM, and Unsloth.

7m read timeFrom blog.dailydoseofds.com
Post cover image
Table of contents
Free Claude Code Course with Lydia Hallie, Anthropic & Master.devSpeculation is all you need!Build Agents that can learn like humans ​
12 Impressions