vLLM
Related tags:
Posts about ai-inferencePosts about vllmPosts about gpuPosts about devopsPosts about mixture-of-expertsPosts about cicd
Optimizing vLLM on Arm CPUsParallel All the Way Down: Beyond Single-Token Generation with Speculative DecodingKimi K3 Is Here: Efficient Day-0 Support on vLLMAnnouncing vLLM AFD Plugin: Disaggregating Attention and FFN for Flexible MoE ServingA Preview of Production-Scale Kimi K3 Support on vLLMKeeping vLLM Production Quality: A Look Inside CI, Benchmarking, and the Release ProcessTML Inkling on vLLM: Day-0 Support with Optimized PerformancevLLM x TileRT: Specialized Decode for Latency-Critical ServingEAGLE-3 Speculative Decoding on AMD Instinct GPUs: Training and Serving with vLLM and AMD QuarkvLLM × HPC-Ops: High-Performance Attention and MoE Backends from Tencent Hunyuan