<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/glm-5-3-didn-t-change-the-base-model-where-did-its-coding-gains-come-from--thx8jk2fo" -->

---
title: GLM-5.3 didn’t change the base model — where did its...
description: Z.ai released GLM-5.3, a coding and agent model that reuses the same base model as GLM-5.2 but gains its performance boost entirely from expanded...
canonical: https://daily.dev/posts/glm-5-3-didn-t-change-the-base-model-where-did-its-coding-gains-come-from--thx8jk2fo
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: GLM-5.3 didn’t change the base model — where did its coding gains come from? | daily.dev
og:description: Z.ai released GLM-5.3, a coding and agent model that reuses the same base model as GLM-5.2 but gains its performance boost entirely from expanded...
og:url: https://daily.dev/posts/glm-5-3-didn-t-change-the-base-model-where-did-its-coding-gains-come-from--thx8jk2fo
og:image: https://api.daily.dev/og/posts/thx8jk2FO.png
og:image:alt: GLM-5.3 didn’t change the base model — where did its coding gains come from?
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# GLM-5.3 didn’t change the base model — where did its coding gains come from?

**[The New Stack](https://daily.dev/sources/newstack)** · 5 min read · 0 upvotes · 0 comments

## Summary

Z.ai released GLM-5.3, a coding and agent model that reuses the same base model as GLM-5.2 but gains its performance boost entirely from expanded post-training, including tenfold more long-horizon task environments. Public benchmarks show large jumps: Terminal-Bench 3.0 rose from 4.6 to 28.3, DeepSWE v1.1 from 46.2 to 66.9, and CyberGym from 77.2% to 84.5%, though exploitation-stage benchmarks (ExploitBench) still trail competitors like Mythos 5 and GPT-5.6 Sol. The model supports a 1-million-token context window and configurable reasoning effort levels, and is already usable via Z.ai's GLM Coding Plan with Claude Code, Cline, OpenCode and Codex, but direct API access and open weights won't arrive for two weeks. A migration quirk redirects Coding Plan calls to GLM-5.2 or GLM-5.1 automatically to GLM-5.3, complicating clean version comparisons.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://thenewstack.io/glm-5-3-post-training-coding>

## Questions this post answers

### Where did GLM-5.3's coding performance gains come from if it uses the same base model as GLM-5.2?

The gains come entirely from expanded post-training rather than a new base model. Z.ai exposed GLM-5.3 to roughly tenfold more long-horizon task environments and broadened its access to developer tools, with some training tasks simulating full software lifecycle work equivalent to a senior engineer's multi-day workload.

_Developers weighing model upgrades can track how post-training scaling reshapes coding benchmarks on daily.dev._

### When will Z.ai release the model weights for GLM-5.3?

Weights are scheduled to arrive about two weeks after the model's release on the GLM Coding Plan, following a hardening and safety-testing period. Until then, GLM-5.3 is only accessible through Z.ai's GLM Coding Plan via Anthropic-compatible and OpenAI-compatible endpoints with tools like Claude Code, Cline, OpenCode, and Codex, not through direct API or local deployment.

_Teams planning to self-host or benchmark new open models can follow release timelines like this on daily.dev._

### How does GLM-5.3 perform on agentic coding benchmarks compared to GLM-5.2?

GLM-5.3 shows large jumps on public agentic coding evaluations: Terminal-Bench 3.0 rose from 4.6 to 28.3, DeepSWE v1.1 climbed from 46.2 to 66.9, and Agents' Last Exam moved from 23.8 to 28.5. Its DeepSWE score of 66.9 lands close to Google's Gemini 3.7 Flash at 65%, though differing test harnesses make direct comparisons unreliable.

_Evaluating which coding model to adopt gets easier when developers compare benchmark shifts like these on daily.dev._

## Similar posts on daily.dev

- [GLM-5.2: Built for Long-Horizon Tasks](https://daily.dev/posts/glm-5-2-built-for-long-horizon-tasks-stmnmwjok) · Hugging Face · 33 upvotes · 4 comments
- [Z.ai pitches GLM-5.2 for long-running software engineering tasks](https://daily.dev/posts/z-ai-pitches-glm-5-2-for-long-running-software-engineering-tasks-3x3e7fy5z) · InfoWorld · 1 upvotes · 0 comments

---

Tags: [#llm](https://daily.dev/tags/llm), [#ai-agents](https://daily.dev/tags/ai-agents)

[View this post on daily.dev](https://daily.dev/posts/glm-5-3-didn-t-change-the-base-model-where-did-its-coding-gains-come-from--thx8jk2fo)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"GLM-5.3 didn’t change the base model — where did its coding gains come from?","url":"https://daily.dev/posts/glm-5-3-didn-t-change-the-base-model-where-did-its-coding-gains-come-from--thx8jk2fo","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/glm-5-3-didn-t-change-the-base-model-where-did-its-coding-gains-come-from--thx8jk2fo"},"datePublished":"2026-08-14T15:27:07.778Z","dateModified":"2026-08-16T05:50:52.003Z","description":"Z.ai released GLM-5.3, a coding and agent model that reuses the same base model as GLM-5.2 but gains its performance boost entirely from expanded...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/f5e8a14fe075e4eea38f5687dbb09ba4?_a=AQAEuop","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/f5e8a14fe075e4eea38f5687dbb09ba4?_a=AQAEuop","isAccessibleForFree":true,"articleSection":"The New Stack","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"The New Stack","logo":"https://media.daily.dev/image/upload/t_logo,f_auto/v1/logos/newstack","url":"https://daily.dev/sources/newstack"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/glm-5-3-didn-t-change-the-base-model-where-did-its-coding-gains-come-from--thx8jk2fo","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":0},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"llm,ai-agents","timeRequired":"PT5M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"The New Stack","item":"https://daily.dev/sources/newstack"},{"@type":"ListItem","position":3,"name":"GLM-5.3 didn’t change the base model — where did its coding gains come from?"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/glm-5-3-didn-t-change-the-base-model-where-did-its-coding-gains-come-from--thx8jk2fo#faq","mainEntity":[{"@type":"Question","name":"Where did GLM-5.3's coding performance gains come from if it uses the same base model as GLM-5.2?","acceptedAnswer":{"@type":"Answer","text":"The gains come entirely from expanded post-training rather than a new base model. Z.ai exposed GLM-5.3 to roughly tenfold more long-horizon task environments and broadened its access to developer tools, with some training tasks simulating full software lifecycle work equivalent to a senior engineer's multi-day workload. Developers weighing model upgrades can track how post-training scaling reshapes coding benchmarks on daily.dev."}},{"@type":"Question","name":"When will Z.ai release the model weights for GLM-5.3?","acceptedAnswer":{"@type":"Answer","text":"Weights are scheduled to arrive about two weeks after the model's release on the GLM Coding Plan, following a hardening and safety-testing period. Until then, GLM-5.3 is only accessible through Z.ai's GLM Coding Plan via Anthropic-compatible and OpenAI-compatible endpoints with tools like Claude Code, Cline, OpenCode, and Codex, not through direct API or local deployment. Teams planning to self-host or benchmark new open models can follow release timelines like this on daily.dev."}},{"@type":"Question","name":"How does GLM-5.3 perform on agentic coding benchmarks compared to GLM-5.2?","acceptedAnswer":{"@type":"Answer","text":"GLM-5.3 shows large jumps on public agentic coding evaluations: Terminal-Bench 3.0 rose from 4.6 to 28.3, DeepSWE v1.1 climbed from 46.2 to 66.9, and Agents' Last Exam moved from 23.8 to 28.5. Its DeepSWE score of 66.9 lands close to Google's Gemini 3.7 Flash at 65%, though differing test harnesses make direct comparisons unreliable. Evaluating which coding model to adopt gets easier when developers compare benchmark shifts like these on daily.dev."}}]}
```

