vLLM
Read post

EAGLE-3 Speculative Decoding on AMD Instinct GPUs: Training and Serving with vLLM and AMD Quark

EAGLE-3 speculative decoding is now fully supported on AMD Instinct GPUs through a vLLM-centric pipeline built by the AMD Quark team. The workflow covers training EAGLE-3 draft models (using vLLM for on-policy data synthesis, hidden-state extraction, and in-loop evaluation), quantizing both target and draft models with AMD Quark MXFP4/FP8, and serving on ROCm/vLLM. Benchmarks on AMD Instinct MI355X show 1.69x–2.00x throughput gains for Kimi-K2.5 and up to 1.79x for MiniMax-M2.5 at 1K/1K workloads. The MiniMax-M3 EAGLE-3 draft trained in this pipeline achieves an average acceptance length of 2.77 across 11 domains on SPEED-Bench, remaining stable from 1K to 32K context. Day-0 MXFP4 checkpoints for major models are published on Hugging Face and run directly on ROCm/vLLM.

    #ai-inference#vllm
Jul 14•12m read time•From vllm.ai
Post cover image
Table of contents
Why Speculative Decoding and EAGLE-3 MatterAMD Quark MXFP4: Day-0 Quantization for Mainstream LLMsTraining EAGLE-3 Draft Models with vLLMOne Team, End to End: The AMD Quark AdvantageAcceleration ResultsSummaryAcknowledgementsAdditional Resources
305 Impressions
vLLM's image
vLLM

76 Followers

•

163 Upvotes

Would you recommend this post?

Copy link
WhatsApp
Facebook
X
New Squad
  • © 2026 Daily Dev Ltd.
  • Guidelines
  • Explore
  • Tags
  • Sources
  • Squads
  • Leaderboard