Under 5 minutes to a deployed LLM endpoint — Audry Hsu, RunPod

This title could be clearer and more informative.Try out Clickbait Shieldfor free (5 uses left this month).

A conference talk introducing RunPod, a cloud GPU infrastructure platform. The speaker demonstrates how to deploy an LLM endpoint in under 5 minutes using RunPod's serverless product and hub listings. The demo covers selecting a pre-configured open-source LLM from the hub, deploying it via the console with vLLM, configuring auto-scaling workers and spending caps, and sending requests to the provisioned HTTP endpoint. Key platform features covered include pods, serverless auto-scaling, multi-node clusters, and the hub repository. RunPod reports 500K+ developers, 30+ global data centers, and $120M ARR.

13m watch time
93 Impressions