IBM Research implemented paged attention (the core kernel of vLLM) in Helion, PyTorch's new domain-specific language for portable high-performance kernels. The Helion implementation required 133 lines versus 295 in Triton, with automatic handling of tiling, masking, and boundaries. On NVIDIA H100, the Helion kernel matched or exceeded Triton performance for decode workloads (132-153% of Triton speed) and was competitive for prefill. End-to-end vLLM benchmarks showed dynamic shapes achieved 96% of Triton throughput on H100, while static shapes caused severe JIT overhead. The team found Helion's autotuner powerful but time-consuming (10 hours for quick mode), and concluded that dynamic shapes are essential for production inference servers with diverse request patterns.

25m read timeFrom pytorch.org
Post cover image
Table of contents
Brief Background to vLLM, Triton, and HelionImplementation Details: How to write Paged Attention in Tiled PyTorchPerformance EvaluationAcknowledgments
177 Impressions