1. /
  2. Sources/
  3. vLLM
vLLM logo

vLLM

Related tags:

#ai-inference#vllm#gpu#devops#mixture-of-experts#cicd
Posts about ai-inferencePosts about vllmPosts about gpuPosts about devopsPosts about mixture-of-expertsPosts about cicd
Optimizing vLLM on Arm CPUsParallel All the Way Down: Beyond Single-Token Generation with Speculative DecodingKimi K3 Is Here: Efficient Day-0 Support on vLLMAnnouncing vLLM AFD Plugin: Disaggregating Attention and FFN for Flexible MoE ServingA Preview of Production-Scale Kimi K3 Support on vLLMKeeping vLLM Production Quality: A Look Inside CI, Benchmarking, and the Release ProcessTML Inkling on vLLM: Day-0 Support with Optimized PerformancevLLM x TileRT: Specialized Decode for Latency-Critical ServingEAGLE-3 Speculative Decoding on AMD Instinct GPUs: Training and Serving with vLLM and AMD QuarkvLLM × HPC-Ops: High-Performance Attention and MoE Backends from Tencent Hunyuan

Most upvoted posts from vLLM

Best discussed posts from vLLM

All posts from vLLM