<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/stepfun-step-5-preview-a-600b-moe-model-worth-trying-for-coding-agents-t2cxiew0c" -->

---
title: StepFun Step 5 Preview: a 600B MoE model worth trying...
description: StepFun's newly released Step 5 Preview is a 600B-parameter mixture-of-experts model (27B active) with a 1M token context window and vision support, available...
canonical: https://daily.dev/posts/stepfun-step-5-preview-a-600b-moe-model-worth-trying-for-coding-agents-t2cxiew0c
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: StepFun Step 5 Preview: a 600B MoE model worth trying for coding agents | daily.dev
og:description: StepFun's newly released Step 5 Preview is a 600B-parameter mixture-of-experts model (27B active) with a 1M token context window and vision support, available...
og:url: https://daily.dev/posts/stepfun-step-5-preview-a-600b-moe-model-worth-trying-for-coding-agents-t2cxiew0c
og:image: https://api.daily.dev/og/posts/t2CXiEw0c.png
og:image:alt: StepFun Step 5 Preview: a 600B MoE model worth trying for coding agents
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# StepFun Step 5 Preview: a 600B MoE model worth trying for coding agents

**[Collections](https://daily.dev/sources/collections)** · 2 min read · 1 upvotes · 0 comments

## Summary

StepFun's newly released Step 5 Preview is a 600B-parameter mixture-of-experts model (27B active) with a 1M token context window and vision support, available via StepFun AI Studio and API. Hands-on testing as a coding agent against GLM 5.3 on two real repo tasks found both models produced correct, regression-free code, but Step 5 Preview reliably recognized when it was done and stopped, while GLM 5.3 kept running until hitting the step limit. In a 368K-token synthetic long-context test with five hidden clues, Step 5 Preview correctly found and connected all five in about 90 seconds. It's recommended for unattended agent runs, bug fixes in unfamiliar codebases, and long-context synthesis tasks.

## Content

StepFun released Step 5 Preview, a mixture-of-experts model with 600B total parameters and 27B active, a 1M token context window, and vision capabilities. It's available now through StepFun AI Studio and the API.

I got early access and spent time testing it as a coding agent. Short version: it's on the Pareto frontier for cost vs. capability, comparable to GLM 5.3 and Kimi K3, and it has one behavior that sets it apart from most models in this class - it knows when to stop.

## How I tested it

I ran Step 5 Preview and GLM 5.3 against two real tasks in the same repo, at the same commit. No follow-up prompts, no retries.

**Task 1:** A bug where numeric filters silently returned zero rows for decimals and negatives.

**Task 2:** A feature requiring new routes, permission gating, and a refactor of the background task supervisor.

Both models completed both tasks correctly. Every held-out test passed, no regressions, and neither model weakened an existing test to get there.

The difference showed up after the work was done. Step 5 Preview finished, checked its work, and declared itself done - both times. GLM 5.3 wrote correct code both times and then kept going until the step limit ended the run.

On the bug fix specifically, Step 5 Preview wrote a shorter patch that matched the approach the Datasette maintainer used in the actual commit. It also added its own tests without being asked.

## Long-context test

Separately, I generated about 368K tokens of fake incident tickets and hid five clues inside them that together explain an outage. When I asked for the root cause, it found all five clues and connected them correctly in about 90 seconds.

## When to reach for it

Based on these tests, Step 5 Preview is a good fit for:

- Unattended agent runs where a clear completion signal matters
- Bug fixes in unfamiliar codebases
- Long-context tasks where you need the model to actually synthesize information rather than skim

It's worth trying in your coding agent setup, especially if runaway tool calls are a problem you've been working around.

*Thanks to the StepFun team for early access.*

## Questions this post answers

### What are the specs of StepFun's Step 5 Preview model?

Step 5 Preview is a mixture-of-experts model with 600 billion total parameters and 27 billion active parameters, a 1 million token context window, and vision capabilities. It is available through StepFun AI Studio and via API, positioned as a coding-agent-capable model comparable in cost-to-capability terms to GLM 5.3 and Kimi K3.

_Developers weighing new coding-agent models can track releases like Step 5 Preview on daily.dev._

### How does Step 5 Preview compare to GLM 5.3 for coding agent tasks?

Both models completed the same two real-repo tasks correctly with no regressions in a head-to-head test, but Step 5 Preview recognized when the work was done and stopped, while GLM 5.3 kept generating output until hitting the step limit. Step 5 Preview also wrote a shorter, more targeted patch and added its own tests unprompted on a bug-fix task.

_For anyone comparing coding agent models on real tasks, daily.dev surfaces evaluations like this one._

### How well does Step 5 Preview handle very long context for reasoning tasks?

In a test with about 368,000 tokens of synthetic incident tickets containing five hidden clues explaining an outage, Step 5 Preview found all five clues and correctly connected them to identify the root cause in roughly 90 seconds. This suggests strong long-context synthesis rather than mere retrieval or skimming.

_Engineers evaluating long-context models for real workloads can follow results like this via daily.dev._

## Community take

How the wider developer community reacted, aggregated from 2 discussions and 21 comments across x (as of 2026-09-22).

**TL;DR:** Reactions to Step 5 Preview are largely enthusiastic, with people highlighting its ability to stop when done and its cost/capability tradeoff, though a few push back on unverified cost claims and question GLM 5.3's reliability.

**Sentiment:** 55% positive · 35% mixed · 10% skeptical

**The case for**

- Knowing when to stop is seen as an underrated but important capability that saves cost on agent runs.
- The model is viewed as a strong point on the cost-vs-capability tradeoff, prompting people to want to try it.
- Retrieving five clues across 368K tokens in 90 seconds impressed several commenters.

**The pushback**

- Some argue the 'pareto frontier' framing is unverified without concrete $/task comparisons against GLM and Kimi on the same eval.
- One commenter says fake/synthetic test tickets make clues too easy to spot compared to real-world noise.
- One commenter dismisses GLM 5.3 entirely as unreliable, questioning its use as a baseline for comparison.

**By community**

- x (positive): Replies mostly praise the model's stopping behavior and cost/capability positioning, with a handful of skeptical notes about unverified cost claims and test realism.

**Hottest debate:** Whether the 'pareto frontier' cost/capability claim is meaningful without concrete $/task benchmarks against rival models.

**Open questions**

- What are the actual $/task costs compared to GLM 5.3 and Kimi K3 on the same benchmark?
- How would the model perform on long-context tasks with more realistic, noisier data rather than synthetic clues?

**Highlights**

> @omarsar0 the pattern is the story: glm 5.3, kimi k3, now step 5 - chinese labs are turning frontier coding into a commodity while the west charges premium. "knows when to stop" is also the one skill half the agents on this app never learn
> — [BriskFalcon\_284 on x · 1 points](https://x.com/BriskFalcon_284/status/2102039711720091691)

> @omarsar0 Pareto frontier for cost vs capability is doing a lot of unverified work until someone posts the actual $/task number next to GLM and Kimi on the same eval set. "Impressive" and "cheap" aren't the same claim.
> — [ukrroot on x](https://x.com/ukrroot/status/2102049660529459636)

> @omarsar0 Fake tickets make the clues stand out. Real noise is harder.
> — [adenshepard on x](https://x.com/adenshepard/status/2102063385302925364)

> @omarsar0 GLM 5.3 is not a worthy contender. It can’t even get a save to PDF feature to work in most browsers. I would drop it altogether from LLM as a judge and I tried to use it a lot!
> — [atilab on x](https://x.com/atilab/status/2102176113019474406)

> @omarsar0 Pulling together 5 clues across 368K tokens in 90 seconds is seriously impressive. Did you have to guide the prompt heavily, or did it reason through the incident tickets zero-shot?
> — [ian\_kim58307 on x · 1 comments](https://x.com/ian_kim58307/status/2102039282655654014)

**Source threads**

- [x](https://x.com/omarsar0/status/2102214865779659208) · 0 points · 0 comments
- [x](https://x.com/omarsar0/status/2102035017262149640) · 1 points · 21 comments

## Similar posts on daily.dev

- [GLM-5.2: Built for Long-Horizon Tasks](https://daily.dev/posts/glm-5-2-built-for-long-horizon-tasks-stmnmwjok) · Hugging Face · 33 upvotes · 4 comments
- [Z.ai pitches GLM-5.2 for long-running software engineering tasks](https://daily.dev/posts/z-ai-pitches-glm-5-2-for-long-running-software-engineering-tasks-3x3e7fy5z) · InfoWorld · 1 upvotes · 0 comments

---

Tags: [#llm](https://daily.dev/tags/llm), [#ai-agents](https://daily.dev/tags/ai-agents), [#mixture-of-experts](https://daily.dev/tags/mixture-of-experts)

[View this post on daily.dev](https://daily.dev/posts/stepfun-step-5-preview-a-600b-moe-model-worth-trying-for-coding-agents-t2cxiew0c)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"StepFun Step 5 Preview: a 600B MoE model worth trying for coding agents","url":"https://daily.dev/posts/stepfun-step-5-preview-a-600b-moe-model-worth-trying-for-coding-agents-t2cxiew0c","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/stepfun-step-5-preview-a-600b-moe-model-worth-trying-for-coding-agents-t2cxiew0c"},"datePublished":"2026-09-22T01:54:35.126Z","dateModified":"2026-09-22T01:55:22.943Z","description":"StepFun's newly released Step 5 Preview is a 600B-parameter mixture-of-experts model (27B active) with a 1M token context window and vision support, available...","isAccessibleForFree":true,"articleSection":"Collections","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Collections","logo":"https://media.daily.dev/image/upload/s--fk_6ycEi--/f_auto,q_auto/v1780996001/logos/collections?_a=BAMAMiWQ0","url":"https://daily.dev/sources/collections"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/stepfun-step-5-preview-a-600b-moe-model-worth-trying-for-coding-agents-t2cxiew0c","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":1},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"llm,ai-agents,mixture-of-experts","timeRequired":"PT2M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Collections","item":"https://daily.dev/sources/collections"},{"@type":"ListItem","position":3,"name":"StepFun Step 5 Preview: a 600B MoE model worth trying for coding agents"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/stepfun-step-5-preview-a-600b-moe-model-worth-trying-for-coding-agents-t2cxiew0c#faq","mainEntity":[{"@type":"Question","name":"What are the specs of StepFun's Step 5 Preview model?","acceptedAnswer":{"@type":"Answer","text":"Step 5 Preview is a mixture-of-experts model with 600 billion total parameters and 27 billion active parameters, a 1 million token context window, and vision capabilities. It is available through StepFun AI Studio and via API, positioned as a coding-agent-capable model comparable in cost-to-capability terms to GLM 5.3 and Kimi K3. Developers weighing new coding-agent models can track releases like Step 5 Preview on daily.dev."}},{"@type":"Question","name":"How does Step 5 Preview compare to GLM 5.3 for coding agent tasks?","acceptedAnswer":{"@type":"Answer","text":"Both models completed the same two real-repo tasks correctly with no regressions in a head-to-head test, but Step 5 Preview recognized when the work was done and stopped, while GLM 5.3 kept generating output until hitting the step limit. Step 5 Preview also wrote a shorter, more targeted patch and added its own tests unprompted on a bug-fix task. For anyone comparing coding agent models on real tasks, daily.dev surfaces evaluations like this one."}},{"@type":"Question","name":"How well does Step 5 Preview handle very long context for reasoning tasks?","acceptedAnswer":{"@type":"Answer","text":"In a test with about 368,000 tokens of synthetic incident tickets containing five hidden clues explaining an outage, Step 5 Preview found all five clues and correctly connected them to identify the root cause in roughly 90 seconds. This suggests strong long-context synthesis rather than mere retrieval or skimming. Engineers evaluating long-context models for real workloads can follow results like this via daily.dev."}}]}
```

