A walkthrough of the Kronk AI project, which uses llama.cpp as its inference engine via a Go API called Isma. The talk covers how Kronk differs from Ollama by staying on the latest llama.cpp versions (avoiding the CGO binding lag that keeps Ollama on older releases), how to lock into specific llama.cpp versions for production stability, and the workflow for handling breaking changes. The presenter also notes that llama.cpp recently joined forces with Hugging Face, improving its funding and development pace.
•3m watch time
1 Impression