<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/running-large-language-models-locally-your-own-chatgpt-like-ai-in-c--cxjpqiwzd" -->

---
title: Running Large Language Models locally – Your own...
description: A walkthrough shows how to run open-source large language models like WizardLM locally on a CPU using the .NET library LLamaSharp, a C# binding for llama.cpp....
canonical: https://daily.dev/posts/running-large-language-models-locally-your-own-chatgpt-like-ai-in-c--cxjpqiwzd
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: Running Large Language Models locally – Your own ChatGPT-like AI in C# | daily.dev
og:description: A walkthrough shows how to run open-source large language models like WizardLM locally on a CPU using the .NET library LLamaSharp, a C# binding for llama.cpp....
og:url: https://daily.dev/posts/running-large-language-models-locally-your-own-chatgpt-like-ai-in-c--cxjpqiwzd
og:image: https://api.daily.dev/og/posts/cXjpQiwzD.png
og:image:alt: Running Large Language Models locally – Your own ChatGPT-like AI in C#
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Running Large Language Models locally – Your own ChatGPT-like AI in C#

**[Maarten Balliauw](https://daily.dev/sources/maartenballiauw)** · 7 min read · 0 upvotes · 0 comments

## Summary

A walkthrough shows how to run open-source large language models like WizardLM locally on a CPU using the .NET library LLamaSharp, a C# binding for llama.cpp. It covers the history of leaked LLaMA weights spawning derivatives like Alpaca and Vicuna, downloading a GGML model, setting up a console app with the LLamaSharp and LLamaSharp.Backend.Cpu packages, and building a simple chat session with custom prompts, including a Homer Simpson persona and a pair-programming assistant example.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://blog.maartenballiauw.be/posts/2023-06-15-running-large-language-models-locally-your-own-chatgpt-like-ai-in-csharp>

## Questions this post answers

### How do I run a ChatGPT-like LLM locally in C# using LLamaSharp?

Install the LLamaSharp and LLamaSharp.Backend.Cpu NuGet packages in a console application, then load a GGML model file such as wizardLM-7B.ggmlv3.q4_1.bin with LLamaModel and LLamaParams, specifying context size and other parameters. Create a ChatSession with a system prompt and an anti-prompt like 'User: ', then call session.Chat() in a loop to stream token-by-token responses, all running on CPU without calling OpenAI.

_Developers experimenting with local LLM setups can track C# and LLM tooling updates on daily.dev._

### Which open-source models can llama.cpp run on a CPU?

Llama.cpp, a C++ implementation by Georgi Gerganov, can run LLaMA and its derivatives including Alpaca, GPT4All, Vicuna, Koala, OpenBuddy, and WizardLM on consumer CPU hardware. It emerged after Meta's LLaMA weights were leaked despite the model being restricted to non-commercial research use, sparking a wave of fine-tuned variants from universities and the open-source community.

_Those comparing local inference engines can follow llama.cpp ecosystem developments on daily.dev._

### How much disk space and memory does the WizardLM 7B GGML model need to run locally?

The WizardLM-7B GGML model variants require between 2.8 GB and 8 GB of disk space depending on quantization, and up to 10 GB of RAM to run inference. The wizardLM-7B.ggmlv3.q4_1.bin variant offers a balance between accuracy and inference speed, which matters when running purely on a CPU rather than a GPU.

_Developers sizing hardware for local LLM experiments can keep tabs on model requirements via daily.dev._

## Similar posts on daily.dev

- [How To Run an Open-Source LLM on Your Personal Computer – Run Ollama Locally](https://daily.dev/posts/how-to-run-an-open-source-llm-on-your-personal-computer-run-ollama-locally-07pat0hqe) · freeCodeCamp · 11 upvotes · 0 comments
- [Run a Local AI Model with Ollama in 15 Minutes](https://daily.dev/posts/run-a-local-ai-model-with-ollama-in-15-minutes-5wx17jt3x) · Machine Learning Mastery · 1 upvotes · 0 comments
- [I Tried This Open Source ChatGPT Alternative on Linux, But Went Back to Ollama](https://daily.dev/posts/i-tried-this-open-source-chatgpt-alternative-on-linux-but-went-back-to-ollama-gbag5xspb) · It's Foss · 0 upvotes · 0 comments
- [How to Use Ollama to Run Large Language Models Locally – Real Python](https://daily.dev/posts/how-to-use-ollama-to-run-large-language-models-locally-real-python-fhurefciw) · Real Python · 3 upvotes · 0 comments

---

Tags: [#llm](https://daily.dev/tags/llm), [#c#](https://daily.dev/tags/c#), [#local-ai](https://daily.dev/tags/local-ai), [#llama-cpp](https://daily.dev/tags/llama-cpp)

[View this post on daily.dev](https://daily.dev/posts/running-large-language-models-locally-your-own-chatgpt-like-ai-in-c--cxjpqiwzd)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"Running Large Language Models locally – Your own ChatGPT-like AI in C#","url":"https://daily.dev/posts/running-large-language-models-locally-your-own-chatgpt-like-ai-in-c--cxjpqiwzd","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/running-large-language-models-locally-your-own-chatgpt-like-ai-in-c--cxjpqiwzd"},"datePublished":"2026-10-04T13:04:14.993Z","dateModified":"2026-10-04T13:05:19.105Z","description":"A walkthrough shows how to run open-source large language models like WizardLM locally on a CPU using the .NET library LLamaSharp, a C# binding for llama.cpp....","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/e0b287694c4b362e77370d906ddb37da?_a=AQAEuop","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/e0b287694c4b362e77370d906ddb37da?_a=AQAEuop","isAccessibleForFree":true,"articleSection":"Maarten Balliauw","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Maarten Balliauw","logo":"https://media.daily.dev/image/upload/s--rOTg3VYO--/f_auto,q_auto/v1791119043/logos/maartenballiauw?_a=BAMAMicg0","url":"https://daily.dev/sources/maartenballiauw"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/running-large-language-models-locally-your-own-chatgpt-like-ai-in-c--cxjpqiwzd","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":0},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"llm,c#,local-ai,llama-cpp","timeRequired":"PT7M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Maarten Balliauw","item":"https://daily.dev/sources/maartenballiauw"},{"@type":"ListItem","position":3,"name":"Running Large Language Models locally – Your own ChatGPT-like AI in C#"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/running-large-language-models-locally-your-own-chatgpt-like-ai-in-c--cxjpqiwzd#faq","mainEntity":[{"@type":"Question","name":"How do I run a ChatGPT-like LLM locally in C# using LLamaSharp?","acceptedAnswer":{"@type":"Answer","text":"Install the LLamaSharp and LLamaSharp.Backend.Cpu NuGet packages in a console application, then load a GGML model file such as wizardLM-7B.ggmlv3.q4_1.bin with LLamaModel and LLamaParams, specifying context size and other parameters. Create a ChatSession with a system prompt and an anti-prompt like 'User: ', then call session.Chat() in a loop to stream token-by-token responses, all running on CPU without calling OpenAI. Developers experimenting with local LLM setups can track C# and LLM tooling updates on daily.dev."}},{"@type":"Question","name":"Which open-source models can llama.cpp run on a CPU?","acceptedAnswer":{"@type":"Answer","text":"Llama.cpp, a C++ implementation by Georgi Gerganov, can run LLaMA and its derivatives including Alpaca, GPT4All, Vicuna, Koala, OpenBuddy, and WizardLM on consumer CPU hardware. It emerged after Meta's LLaMA weights were leaked despite the model being restricted to non-commercial research use, sparking a wave of fine-tuned variants from universities and the open-source community. Those comparing local inference engines can follow llama.cpp ecosystem developments on daily.dev."}},{"@type":"Question","name":"How much disk space and memory does the WizardLM 7B GGML model need to run locally?","acceptedAnswer":{"@type":"Answer","text":"The WizardLM-7B GGML model variants require between 2.8 GB and 8 GB of disk space depending on quantization, and up to 10 GB of RAM to run inference. The wizardLM-7B.ggmlv3.q4_1.bin variant offers a balance between accuracy and inference speed, which matters when running purely on a CPU rather than a GPU. Developers sizing hardware for local LLM experiments can keep tabs on model requirements via daily.dev."}}]}
```

