A developer shares their experience migrating from cloud LLMs to locally hosted models running on a Proxmox LXC container with GPU passthrough. After starting with Ollama, they switched to llama.cpp for better performance and efficiency. The setup supports models like Qwen3 and Gemma4 for coding, RAG analysis, and automation pipelines. Open WebUI provides a ChatGPT-like interface with integrations including SearXNG for web access and ComfyUI for image workflows. The author finds the performance trade-off acceptable given the privacy and cost benefits.
Table of contents
Proxmox LXCs are incredible for hosting llama.cppCertain local models have terrific reasoning capabilitiesThe llama-server web UI is pretty neat for my inference tasks1.4K Impressions