Ultimate Go Software Design LIVE: Ep.80
This title could be clearer and more informative.Try out Clickbait Shieldfor free (5 uses left this month).
A live coding session focused on Go-based LLM inference server development (Kron/Cron), covering performance optimizations including a 3x session-to-slot ratio improvement for better CPU/GPU pipeline parallelism, benchmark running methodology, and memory profiling with pprof on image generation models (Stable Diffusion, Flux). The session also explores running small quantized models (4B parameter Qwen) under 16GB VRAM for coding agent use cases, discusses paged attention vs KV cache tradeoffs, and reviews a proposed streaming audio transcription API design.
•1h 58m watch time
130 Impressions