<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/this-ai-has-320-billion-parameters-it-barely-uses-them--qionyfhmr" -->

---
title: This AI Has 320 Billion Parameters. It Barely Uses Them.
description: GLM 5.3 and its smaller sibling GLM 5.3 Flash are new open-weight AI models that pack 320 billion parameters but activate only about 5% per token, using linear...
canonical: https://daily.dev/posts/this-ai-has-320-billion-parameters-it-barely-uses-them--qionyfhmr
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: This AI Has 320 Billion Parameters. It Barely Uses Them. | daily.dev
og:description: GLM 5.3 and its smaller sibling GLM 5.3 Flash are new open-weight AI models that pack 320 billion parameters but activate only about 5% per token, using linear...
og:url: https://daily.dev/posts/this-ai-has-320-billion-parameters-it-barely-uses-them--qionyfhmr
og:image: https://api.daily.dev/og/posts/qIONyfhmr.png
og:image:alt: This AI Has 320 Billion Parameters. It Barely Uses Them.
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# This AI Has 320 Billion Parameters. It Barely Uses Them.

**[Two Minute Papers](https://daily.dev/sources/twoninutepapers)** · 5 min read · 3 upvotes · 0 comments

## Summary

GLM 5.3 and its smaller sibling GLM 5.3 Flash are new open-weight AI models that pack 320 billion parameters but activate only about 5% per token, using linear attention plus sparse attention to summarize nearby context cheaply, and a technique called 'index pool' to compress stored context for faster long-session lookback. The layer count was roughly halved from 92. Flash quickly overtook DeepSeek in usage after being stealthily released under another name, and both models reportedly approach frontier-level performance on some benchmarks when given extended reasoning time, though still requiring thousands of dollars of hardware to run well.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://www.youtube.com/watch?v=w9RDunJACkc>

## Questions this post answers

### What is GLM 5.3 and how does it reduce compute compared to older large language models?

GLM 5.3 is an open-weight large language model with 320 billion parameters, of which about 95% are not activated per token, using a sparse mixture-of-experts design. Its layer count was roughly halved from a prior 92 layers, and it combines linear attention with sparse attention to summarize nearby context cheaply, plus an index-pool technique that compresses stored context so long sessions stay fast and use less memory.

_Track efficiency breakthroughs like GLM 5.3's sparse attention design as they reshape which models developers can run locally, on daily.dev._

### Why did GLM 5.3 Flash overtake DeepSeek in usage after its release?

GLM 5.3 Flash was initially released stealthily under a different name before its identity became known, and it quickly gained more usage than DeepSeek. Both GLM 5.3 and its Flash variant, when allowed to reason longer, get close to frontier-level results on some benchmarks, despite still requiring thousands of dollars of hardware for full performance.

_Developers weighing open-weight models against DeepSeek can follow benchmark shifts like this on daily.dev._

---

Tags: [#open-source](https://daily.dev/tags/open-source), [#llm](https://daily.dev/tags/llm), [#ai-inference](https://daily.dev/tags/ai-inference)

[View this post on daily.dev](https://daily.dev/posts/this-ai-has-320-billion-parameters-it-barely-uses-them--qionyfhmr)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"This AI Has 320 Billion Parameters. It Barely Uses Them.","url":"https://daily.dev/posts/this-ai-has-320-billion-parameters-it-barely-uses-them--qionyfhmr","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/this-ai-has-320-billion-parameters-it-barely-uses-them--qionyfhmr"},"datePublished":"2026-09-01T09:06:51.007Z","dateModified":"2026-09-04T16:03:24.751Z","description":"GLM 5.3 and its smaller sibling GLM 5.3 Flash are new open-weight AI models that pack 320 billion parameters but activate only about 5% per token, using linear...","image":"https://i.ytimg.com/vi/w9RDunJACkc/sddefault.jpg","thumbnailUrl":"https://i.ytimg.com/vi/w9RDunJACkc/sddefault.jpg","isAccessibleForFree":true,"articleSection":"Two Minute Papers","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Two Minute Papers","logo":"https://media.daily.dev/image/upload/s--8KXsSJ5q--/f_auto/v1711188923/logos/twoninutepapers","url":"https://daily.dev/sources/twoninutepapers"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/this-ai-has-320-billion-parameters-it-barely-uses-them--qionyfhmr","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":3},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"open-source,llm,ai-inference","timeRequired":"PT5M","video":{"@type":"VideoObject","name":"This AI Has 320 Billion Parameters. It Barely Uses Them.","description":"GLM 5.3 and its smaller sibling GLM 5.3 Flash are new open-weight AI models that pack 320 billion parameters but activate only about 5% per token, using linear...","thumbnailUrl":"https://i.ytimg.com/vi/w9RDunJACkc/sddefault.jpg","uploadDate":"2026-09-01T09:06:51.007Z","duration":"PT5M","url":"https://api.daily.dev/r/qIONyfhmr","embedUrl":"https://www.youtube.com/embed/w9RDunJACkc"}}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Two Minute Papers","item":"https://daily.dev/sources/twoninutepapers"},{"@type":"ListItem","position":3,"name":"This AI Has 320 Billion Parameters. It Barely Uses Them."}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/this-ai-has-320-billion-parameters-it-barely-uses-them--qionyfhmr#faq","mainEntity":[{"@type":"Question","name":"What is GLM 5.3 and how does it reduce compute compared to older large language models?","acceptedAnswer":{"@type":"Answer","text":"GLM 5.3 is an open-weight large language model with 320 billion parameters, of which about 95% are not activated per token, using a sparse mixture-of-experts design. Its layer count was roughly halved from a prior 92 layers, and it combines linear attention with sparse attention to summarize nearby context cheaply, plus an index-pool technique that compresses stored context so long sessions stay fast and use less memory. Track efficiency breakthroughs like GLM 5.3's sparse attention design as they reshape which models developers can run locally, on daily.dev."}},{"@type":"Question","name":"Why did GLM 5.3 Flash overtake DeepSeek in usage after its release?","acceptedAnswer":{"@type":"Answer","text":"GLM 5.3 Flash was initially released stealthily under a different name before its identity became known, and it quickly gained more usage than DeepSeek. Both GLM 5.3 and its Flash variant, when allowed to reason longer, get close to frontier-level results on some benchmarks, despite still requiring thousands of dollars of hardware for full performance. Developers weighing open-weight models against DeepSeek can follow benchmark shifts like this on daily.dev."}}]}
```

