<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/why-deepseek-is-cheap-at-scale-but-expensive-to-run-locally-0l7watfec" -->

---
title: Why DeepSeek is cheap at scale but expensive to run locally
description: DeepSeek-V3 and similar mixture-of-experts models require high batch sizes to run efficiently due to their architecture. GPUs perform best with large matrix...
canonical: https://daily.dev/posts/why-deepseek-is-cheap-at-scale-but-expensive-to-run-locally-0l7watfec
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: Why DeepSeek is cheap at scale but expensive to run locally | daily.dev
og:description: DeepSeek-V3 and similar mixture-of-experts models require high batch sizes to run efficiently due to their architecture. GPUs perform best with large matrix...
og:url: https://daily.dev/posts/why-deepseek-is-cheap-at-scale-but-expensive-to-run-locally-0l7watfec
og:image: https://api.daily.dev/og/posts/0L7wATfec.png
og:image:alt: Why DeepSeek is cheap at scale but expensive to run locally
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Why DeepSeek is cheap at scale but expensive to run locally

**[Hacker News](https://daily.dev/sources/hn)** · 12 min read · 3 upvotes · 0 comments

## Summary

DeepSeek-V3 and similar mixture-of-experts models require high batch sizes to run efficiently due to their architecture. GPUs perform best with large matrix multiplications, so inference servers batch multiple user requests together. Models with many experts and layers need larger batches to avoid pipeline bubbles and keep all experts busy, which increases latency but dramatically improves throughput. This explains why DeepSeek is cost-effective at scale but inefficient for single-user local deployment.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://www.seangoedecke.com/inference-batching-and-deepseek/>

---

Tags: [#ai](https://daily.dev/tags/ai), [#machine-learning](https://daily.dev/tags/machine-learning), [#performance](https://daily.dev/tags/performance), [#gpu](https://daily.dev/tags/gpu), [#deepseek](https://daily.dev/tags/deepseek)

[View this post on daily.dev](https://daily.dev/posts/why-deepseek-is-cheap-at-scale-but-expensive-to-run-locally-0l7watfec)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"Why DeepSeek is cheap at scale but expensive to run locally","url":"https://daily.dev/posts/why-deepseek-is-cheap-at-scale-but-expensive-to-run-locally-0l7watfec","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/why-deepseek-is-cheap-at-scale-but-expensive-to-run-locally-0l7watfec"},"datePublished":"2025-06-01T14:24:58.845Z","dateModified":"2025-06-01T14:25:24.294Z","description":"DeepSeek-V3 and similar mixture-of-experts models require high batch sizes to run efficiently due to their architecture. GPUs perform best with large matrix...","image":"https://media.daily.dev/image/upload/s--2-1xRawN--/f_auto/v1722860399/public/Placeholder%2011","thumbnailUrl":"https://media.daily.dev/image/upload/s--2-1xRawN--/f_auto/v1722860399/public/Placeholder%2011","isAccessibleForFree":true,"articleSection":"Hacker News","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Hacker News","logo":"https://media.daily.dev/image/upload/t_logo,f_auto/v1/logos/hn","url":"https://daily.dev/sources/hn"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/why-deepseek-is-cheap-at-scale-but-expensive-to-run-locally-0l7watfec","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":3},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"ai,machine-learning,performance,gpu,deepseek","timeRequired":"PT12M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Hacker News","item":"https://daily.dev/sources/hn"},{"@type":"ListItem","position":3,"name":"Why DeepSeek is cheap at scale but expensive to run locally"}]}
```

