metir
metir
Docs
Download on App StoreGet it on Google PlayLoginSign Up
Back to Blog
StepFun
Open Weights
LLM Benchmarks
China AI
AI Models

StepFun Step 5 Preview: A 600B Open-Weight Bet on Agents

StepFun's Step 5 Preview is a 600B sparse MoE with a 1M-token context, frontier-adjacent scores and a promised open-weight release. Here is what the numbers show, and what the caveats mean.

Metir AI TeamSeptember 21, 20269 min read
StepFun Step 5 Preview: A 600B Open-Weight Bet on Agents

On September 20, 2026, the Shanghai lab StepFun released Step 5 Preview, a large sparse mixture-of-experts model built for long-horizon agentic work. The launch is worth reading carefully because it packs three separate claims into one announcement: a specific architecture, a set of frontier-adjacent benchmark scores, and a promise to publish the weights on October 15. Each of those carries a different weight of evidence, and the honest way to read the release is to keep them apart.

Qwen logoQwen
DeepSeek logoDeepSeek
Moonshot AI logoMoonshot AI
Z.ai logoZ.ai
Step 5 Preview lands in a crowded field of Chinese open-weight labs releasing frontier-adjacent models at aggressive prices.

The architecture, stated plainly

Step 5 Preview holds about 600 billion total parameters and activates roughly 27 billion of them per token. That ratio, under five percent active, is the entire point of a sparse mixture-of-experts design: the model can carry a large store of knowledge in its full weight set while only paying to compute a small slice on any given token. It takes a one-million-token context window, produces up to 64,000 tokens of output, and accepts text, images and video as input.

600BTotal parameters
27BActive per token
1M tokensContext window
Oct 15, 2026Promised open-weight date

The pricing on StepFun's own API is the part most likely to shape real decisions: one dollar per million input tokens, five cents per million on a cache hit, and two dollars seventy per million output tokens, with reasoning tokens counted in the output. That is a fraction of what the leading Western closed models charge, which is the recurring story of the Chinese open-weight labs this year and the reason they keep pressure on everyone else's prices.

The benchmarks are strong, adjacent to the frontier, and vendor-reported

StepFun published a comparison table putting Step 5 Preview against models including GPT-6 Astra, Claude Opus 5, Kimi K3 and GLM 5.3. On Artificial Analysis's Intelligence Index, StepFun reports a score of 44, which places the model among the stronger open-weight entries and, by StepFun's account, near much larger models such as Kimi K3 Max.

Close on knowledge tests, a step behind on agentic coding

Step 5 Preview versus GPT-6 Astra and Claude Opus 5 on four evaluations, from StepFun's own launch table. The gap is small on GPQA Diamond and widens on the long-horizon coding benchmarks. Higher is better.

Vendor-reported figures. StepFun runs Step 5 Preview in High mode against rivals' Max modes, so the rows favour the closed models less than a strict match would.

The shape of the results is consistent and informative. On a broad knowledge test like GPQA Diamond, Step 5 Preview sits within a couple of points of the closed frontier. On the long-horizon coding and terminal benchmarks, the gap widens: it trails GPT-6 Astra and Claude Opus 5 on DeepSWE and SWE-Marathon, and falls further behind on Terminal-Bench v4, where StepFun reports 33.3 against Astra's 57.9. That is the pattern you would expect from a very capable model that has not fully closed the distance on the hardest agentic tasks, which are exactly the tasks the model is marketed for.

Two caveats belong next to those numbers, and StepFun states both. First, these are the maker's own figures, not an independent evaluation. Vendor benchmarks are best read as a claim about direction and a description of what a lab optimized for, not a settled ranking. Second, the comparison is not strictly like-for-like: StepFun runs Step 5 Preview in its High-effort setting against rivals' Max settings, and notes that at least one row uses a reduced test set that is "not directly comparable." None of that makes the results fake. It means the true gap to the closed frontier is probably a little wider than the table's framing suggests.

“

A model that is close on knowledge tests and a step behind on agentic coding is exactly what the marketing would not tell you, and exactly what the benchmark shape reveals.

On reading a launch table

Where the parameter budget argument actually lands

The more interesting technical claim is efficiency. If Step 5 Preview reaches an Intelligence Index near 44 with 600 billion total parameters, and a model like Kimi K3 Max reaches a similar score at a far larger total footprint, then StepFun's design is doing more with less on paper.

The same index score at a fraction of the parameter budget

Two open-weight models near the same Artificial Analysis Intelligence Index score, sized by total parameters. The bar length is the total parameter count; the badge is the index score.

Step 5 Preview600B total · 27B active
Index 44
Kimi K3 Max~2.8T total · mixture-of-experts
Index 44

A sparse mixture-of-experts activates only a slice of its weights per token. StepFun's pitch is that a smaller total footprint can reach the same measured intelligence, which matters most to whoever pays to host the weights.

That matters to a specific audience: whoever pays to host the weights. A smaller total parameter count is cheaper to store, shard and serve, and a low active-parameter ratio keeps per-token compute down. For a hosted API that difference shows up in the price you already saw. For an open-weight release, it shows up in how many GPUs a team needs to run the thing at all. The efficiency claim is not a benchmark score; it is an operating-cost argument, and it is the argument most relevant to anyone considering self-hosting.

Close-up of NVIDIA H100 GPU accelerator modules installed in a server chassis
NVIDIA H100 accelerator modules in a server. Open weights only become useful to a team that can serve them, so a smaller total parameter count and a low active-parameter ratio are as much a cost argument as a capability one. Illustrative hardware photo, not StepFun's own infrastructure. Photo by geekerwan, via Wikimedia Commons, CC BY 3.0.

"Preview" and "open weights coming October 15" are doing real work

The word "preview" is not decoration. Right now Step 5 Preview is reachable only through StepFun's own API; it is not yet a downloadable model. StepFun says the open weights will land on October 15, 2026. Until they do, this is an open-weight promise, not an open-weight release, and the two are not the same thing.

The distinction is more than pedantic. An open-weight release is verifiable in a way an API is not. When the weights are public, independent evaluators can reproduce the benchmarks, security researchers can probe the model directly, and teams can fine-tune and run it on their own hardware without a rate limit or a terms-of-service change standing between them and the model. A preview API delivers none of that; it delivers a demo and a price. The claims in the launch table become checkable on the day the weights ship, and not before. That is the date to watch.

What it changes, without the hype

Step through the noise and the release fits a trend that has defined 2026: capable models are arriving from Chinese labs on a near-monthly cadence, priced aggressively, and increasingly shipped with weights that anyone can run. Each one narrows the practical gap for the workloads that do not strictly need the very top of the closed frontier, and each one keeps downward pressure on the price of inference everywhere.

It also sharpens a question every team building on top of these models now faces. When a new frontier-adjacent model lands every few weeks, at a fraction of last quarter's price, and when the rankings reshuffle with each release, committing an entire product to one provider starts to look less like a strategy and more like a bet that this month's leader will still lead next month. The more durable posture is to stay able to move: treat the model as a swappable component, route each task to whatever is strongest and cheapest for it right now, and re-evaluate as the releases keep coming. Platforms like Metir that keep work model-agnostic across providers exist for exactly this reason, because a strong launch in September is not a reason to rebuild a stack around a single vendor.

The honest summary of Step 5 Preview is that it is a serious, well-targeted model that reaches the neighborhood of the closed frontier on knowledge tasks, trails it on the hardest agentic ones, and does so on figures that are largely its maker's own. The open-weight release, if it arrives on schedule and matches the preview, is what would turn a promising launch into a verifiable one. That is the part worth waiting for.

Sources:

  • StepFun Launches Step 5 Preview: A 600B-Total, 27B-Active MoE Model | MarkTechPost
  • Step 5 Preview: Specs, Price, Benchmarks, and the Gaps | CellCog
  • StepFun launches Step 5 Preview with 600B parameters, 1M context, and open weights coming October 15 | DataStudios
  • Step 5 Preview - Intelligence, Performance & Price Analysis | Artificial Analysis
  • StepFun launches a 600B agent model at $1 per million input tokens | RuntimeWire

Image credits

Hero image: the Pudong skyline in Shanghai, China, where StepFun is headquartered, photographed by Ermell, via Wikimedia Commons, released under the CC0 1.0 public domain dedication. The photograph shows the city, not StepFun's offices or the Step 5 model. In-body photograph: a close-up of NVIDIA H100 GPU modules by geekerwan, via Wikimedia Commons, licensed under CC BY 3.0. It is an illustration of AI accelerator hardware and does not depict StepFun's own infrastructure.

Ready to experience AI that adapts to you?

metir brings together the world's best AI models in one seamless experience. Start for free today.

Get Started Free
metir

Agentic Operating System for Professionals buried in meetings, emails and docs.

© 2026 metir. All rights reserved.

Product

  • Features
  • Pricing
  • Research
  • Docs
  • Blog
  • Enterprise

Company

  • Docs
  • Support
  • Careers

Legal

  • Terms of Service
  • Privacy Policy

Personalisation is powerful. Privacy is non-negotiable.

Status: All systems operational