Quick Answer:
Step 5 Preview is StepFun's new flagship, launched on 20/09/2026: a sparse Mixture-of-Experts model with 600 billion total and 27 billion active parameters, a 1 million-token context window and API pricing of $1.00 input and $2.70 output per million tokens. StepFun says it matches the leading open models at about 65% lower cost, and has promised open weights on 15/10/2026. Today it is API-only, every benchmark is self-reported and there is no technical report, so treat it as promising but unproven.
Most model launches claim to be the best. StepFun's claim is more modest and, for that reason, more interesting: Step 5 Preview aims to be the best model you can get at a given price, not the best model full stop. If the numbers hold, that is exactly the trade-off agent builders care about.
This guide pulls together StepFun's own launch material as reported by trade and independent sites, plus the Artificial Analysis index figures those reports cite, and the creator coverage that has pushed Step 5 into the Hermes Agent community. We flag every figure as vendor-reported, because that is what it is. We have not run Step 5 ourselves.
Julian Goldie introduces Step 5 and pitches it as a free Chinese model for agent builders.
Executive Summary
What it is: Step 5 Preview (model ID step-5-preview) is a sparse Mixture-of-Experts reasoning model from StepFun, a Shanghai lab that has shipped the "Step" family across text, vision and audio for a couple of years. It is described by StepFun as a frontier model for complex work and agentic tasks, with a lean towards software engineering and finance.
- Specs: 600B total parameters, 27B active per token (about 4.5%), 1M-token context, up to 64k output tokens, text output only, image input (one review also reports short video clips).
- Price: $1.00 input (cache miss), $0.05 input (cache hit), $2.70 output per million tokens, as listed on StepFun's pricing page and echoed by Artificial Analysis.
- Independent signal: Artificial Analysis assigns an Intelligence Index of 44, which one review describes as roughly level with GLM-5.3 and Kimi K3.
- Open weights: promised for 15/10/2026. As of 10/10/2026 there is no model card, licence or weights.
- Self-reported benchmarks: DeepSWE v1.1 67.7, StepCodeBench 49.0, ProgramBench 80.5, Terminal-Bench v4 33.3, FrontierFinance 66.4.
- Caveats: no technical report, StepCodeBench is StepFun's own benchmark, early community reactions flag verbosity, and the 27B active count is heavier than some rival "flash" tiers.
Our view: Step 5 Preview is a credible bid for the "best value open tier" position, and the cache-hit price is genuinely attractive for agent loops. It is a preview with unaudited numbers, so the sensible move is to test it on your own tasks and wait for the weights and independent evaluations before committing.
Who StepFun Is and Why Step 5 Matters
StepFun (Chinese name 阶跃星辰) is one of the Chinese labs that has combined closed API products with a healthy open-weights footprint on Hugging Face. Its smaller Flash line shipped under Apache 2.0, which is why the market is hopeful, though not certain, about the Step 5 licence. Step 5 is the lab's attempt to move from capable mid-tier models to a contender in the broader frontier conversation that already includes GLM-5.3, Kimi K3, Qwen 3.8 Max and the DeepSeek V4 family.
It also arrives in the same fortnight as Reflection AI's Beam, a 501B open-weight model from the United States. Together they show where the open-model race is heading: very large sparse models with small active footprints, reinforcement learning at scale, and an argument about cost per solved task rather than parameter count.
Architecture: 600B Total, 27B Active
In a dense model every parameter runs for every token. In a sparse Mixture-of-Experts model a router chooses a small subset of expert sub-networks for each token, so only a slice of the network fires. Step 5 holds 600 billion parameters but activates 27 billion per token, which is 4.5% of the total and means about 95.5% sit idle for any one token. That is the economic logic behind the pricing: StepFun can charge like a mid-tier model while drawing on a much larger pool of specialised capacity.
| Specification | Step 5 Preview |
|---|---|
| Model ID | step-5-preview |
| Architecture | Sparse Mixture-of-Experts |
| Total / active parameters | 600B / 27B |
| Context window | 1M tokens |
| Max output | 64k tokens |
| Input | Text and images (one review also lists short video clips) |
| Output | Text |
| Reasoning control | reasoning_effort: low, medium or high |
| API availability | From 20/09/2026 |
| Open weights | Promised for 15/10/2026 |
The 27B figure needs careful reading. It does not make Step 5 behave like a dense 27B model, because the larger parameter pool gives the router access to many more specialised experts. It also does not make deployment equivalent to hosting a 27B model: the full 600B weight set still has to be stored and distributed across hardware, so memory, interconnect bandwidth and routing efficiency stay important once the weights ship. A sceptical Hacker News commenter quoted in one review pointed out that 27B active is heavier than some rival "flash" models, so the efficiency win depends on intelligence per active parameter rather than the active count itself. Independent inference benchmarks will settle that, not launch charts.
Capabilities Deep Dive
The 1M-token context
StepFun positions the million-token window around workloads that accumulate state over time: repositories, specifications, issue histories, logs, test output and previous attempts for software agents; filings and reports for financial research. A representative agent loop runs from objective to context inspection, plan, tool call, result, state update, next action and verification, and context grows with every cycle. A bigger window gives the loop room before anything has to be summarised or moved into an external memory system.
Long context is not a free lunch. Filling a window increases prefill work and can hurt latency and cost when an application repeatedly sends hundreds of thousands of tokens, and a long window does not guarantee the agent finds relevant information or recovers from failed steps. For production use you still need retrieval, caching and context selection, which is where the 95% cache-hit discount below becomes relevant.
Reasoning effort and modalities
Step 5 is a reasoning model with a reasoning_effort control set to low, medium or high. Because output tokens are billed inclusive of the reasoning trace, effort is a direct cost dial: higher effort means a longer internal trace and a bigger bill. On the multimodal side, one review says you can pass up to 60 images per request; output is text only. Reports differ on whether video is supported (one lists text, images and video, others list text and images), so check the platform documentation before designing around it.
Coding and finance focus
StepFun leans on software engineering, saying Step 5 can go from software engineering and visual applications to programmable hardware, and that in its expert evaluations about 70% of participants judged it capable of autonomously solving coding tasks of moderately high complexity. That is a survey-style claim from the vendor, not a reproducible benchmark. There is also a Claude Code integration through StepFun's "Step Plan" that exposes the full 1M context, according to the same review, which is a smart way to meet developers where they already work. If you are choosing tooling, our guide to the best AI coding agents compares the harnesses.
Finance is the other stated focus, and it is where the table below shows Step 5's most distinctive result.
Benchmarks: What StepFun Reports
The Pareto-frontier claim
StepFun's central argument is a chart. It plots Artificial Analysis Intelligence Index against cost per task on a log scale, and draws a new Pareto frontier that Step 5 Preview sits on. The vendor's claim is a 65% lower cost than GLM-5.3 and Kimi K3 at a comparable level of intelligence.

Read it as the pitch, not the verdict. The vertical axis is an independent index, which gives the chart some credibility, but the framing and the cost axis are StepFun's, and the cost depends on how many tokens each model spends. The independent figure is the Intelligence Index of 44 and the $1.00 and $2.70 prices that Artificial Analysis lists. One review describes 44 as roughly matching GLM-5.3 and Kimi K3.
The comparison table
| Benchmark | Step 5 Preview | Kimi K3 | GLM-5.3 | GPT-6 Astra | Claude Opus 5 |
|---|---|---|---|---|---|
| DeepSWE v1.1 | 67.7 | 67.5 | 66.9 | 74.1 | 74.0 |
| StepCodeBench | 49.0 | 43.9 | 40.2 | 61.0 | 63.9 |
| ProgramBench | 80.5 | 77.8 | 72.0 | 85.4 | 82.3 |
| Terminal-Bench v4 | 33.3 | 12.6 | 41.9 | 57.9 | 52.3 |
| FrontierFinance | 66.4 | 62.6 | 64.1 | 55.0 | 69.7 |
All numbers above come from StepFun's product page as reproduced by a review. The picture is consistent: at parity with the leading open models, behind the top closed models on raw coding quality, and cheaper. Finance is the standout, where Step 5 edges both open rivals and beats GPT-6 Astra (66.4 against 55.0), though it still trails Claude Opus 5 (69.7). The one clear loss to an open rival is Terminal-Bench v4, where GLM-5.3 scores 41.9 against Step 5's 33.3, though Step 5 is well ahead of Kimi K3's 12.6 on that row. For context on the closed models, see our GPT-6 Astra review and Claude Opus 5 coverage.

Why to be cautious
- Self-reported. There is no technical report or paper yet.
- Home-field benchmark. StepCodeBench is StepFun's own test, said to cover 553 repositories in 33 languages. Models usually look good on tests their own team designed.
- Mixed independent signals. An Intelligence Index of 44 is competitive, but it describes a model in the same band as its open rivals, not above them.
- Community scepticism. Quoted Hacker News reactions include complaints that the demo traces gave away mistakes and that outputs can be padded and verbose. These are anecdotes, not measurements.
Pricing and Access
| Per 1M tokens | Input (cache miss) | Input (cache hit) | Output |
|---|---|---|---|
| step-5-preview | $1.00 (about £0.75) | $0.05 (about £0.04) | $2.70 (about £2.03) |
| step-3.7-flash (for context) | $0.20 | $0.04 | $1.15 |
The 95% cache-hit discount is the standout. Agentic workloads re-send a large, mostly static system prompt on every step, so cached input at about $0.05 per million tokens changes the economics of long loops. Output at $2.70 covers both reasoning trace and answer, so a chatty high-effort run costs real money. There is a free V0 tier for testing, with 5 concurrent requests and 10 requests per minute, and rate limits scale with top-ups. The price compares well with closed frontier models; for the other cheap options see our reviews of DeepSeek V4.1 Flash and Solar Mini 4. Sterling figures are approximate conversions at roughly £0.75 to the dollar.
The Open-Weights Countdown
The timeline matters. StepFun opened paid API access on 20/09/2026 and put a public countdown on its homepage for open weights on 15/10/2026. On launch the Hugging Face repository existed but was an empty placeholder with no weights, model card or licence, according to the reviews we read, and we found no post-release coverage as of 10/10/2026. StepFun's Flash line used Apache 2.0, so a permissive release is likely, but that is an inference. If you are planning on-premises deployment, do not build around weights that have not landed; the date, licence and model card are what to check on 15/10/2026.
The open-weights promise is also what makes Step 5 comparable to Beam. Both are sparse models of roughly 500 to 600 billion parameters, with active counts in the mid-twenties of billions, both with release dates still ahead of them.
Using Step 5 with Agents like Hermes
Step 5 has been picked up quickly by the open agent community. Julian Goldie's video pairing it with Hermes Agent, embedded below, hints at why: an agent harness plus a cheap long-context model is exactly the combination that agent builders want. For background on the harness itself, read our Hermes Agent explainer and the earlier Solar Mini 4 and Hermes piece, which covers another cheap model suited to that setup.
Julian Goldie pairs Step 5 with Hermes Agent and pitches the free access.
Two cautions. First, "free" here means the limited V0 tier and promotional access, not unlimited production use. Second, the more context an agent accumulates, the more it costs, so lean on the cache-hit price and keep a retrieval layer. Reasoning-heavy runs at high effort can burn output tokens quickly.
Real-World Reception Versus Benchmarks
Early reactions follow a familiar pattern. Supporters like the honest positioning, with one commenter noting that StepFun frames itself as the best of the cheaper tier rather than claiming to top the charts. Critics describe earlier Step models as reasoning too much and too long, and worry about waffly output. Both can be true: a model can win on price per unit of intelligence and still be verbose enough to erode the saving. The WorldofAI news roundup embedded below places Step 5 among the week's launches, and is useful for seeing how it is being framed against the leaks and rumours around it.
WorldofAI's weekly roundup, including its take on the Step 5 release.
The practical way to cut through the noise is to run your own evaluation: take 20 to 50 real tasks from your workload, run them at low, medium and high reasoning effort, and record cost, latency and pass rate. That tells you more than any vendor chart, and it is cheap at these prices.
How It Compares
- Versus Kimi K3 and GLM-5.3: similar intelligence, lower price, per StepFun and Artificial Analysis. Step 5 loses to GLM-5.3 on Terminal-Bench v4. See Kimi K3 and GLM-5.3.
- Versus Reflection Beam: Beam is open-weight first and more explicit about its training recipe; Step 5 is API-first and already priced. See our Beam review.
- Versus GPT-6 Astra and Claude Opus 5: behind on raw coding (DeepSWE 67.7 against 74.1 and 74.0), ahead of Astra on FrontierFinance, and far cheaper.
- Versus DeepSeek V4.1 Flash: a heavier active count (27B) for a higher-tier claim; see our DeepSeek V4.1 Flash review for the efficiency-first alternative.
Limitations
- Preview status. Behaviour, pricing and limits may change before general availability.
- No technical report, model card or confirmed licence at the time of writing.
- Self-graded results, including an in-house benchmark.
- Verbosity and reasoning cost. Output tokens include the trace, so long reasoning raises bills.
- Heavy to self-host. 600B parameters need serious hardware even with 27B active.
- Safety evidence. We found no published system card or dangerous-capability evaluation.
- Data governance. It is a hosted API from a Chinese lab; regulated buyers should check data-handling terms before sending sensitive material.
Who Should Use It
Try it if you run agent loops with large, repeated prompts and want the cache-hit price, build coding or finance research tools over big document sets, or want to prepare for an open-weight release you can self-host. Wait if you need a settled model with a model card and independent evaluations, have strict data-residency requirements, or need the strongest possible coding model, where closed frontier models still lead. Test before trusting in every case: the benchmarks are the vendor's.
The Bottom Line
Step 5 Preview is a sensible, honestly positioned launch: a large sparse model that tries to match the best open models at a much lower price, with a pricing structure that suits agents. The evidence so far is a self-reported table, an independent index score that puts it in the same band as its rivals, and a set of community reactions that are positive but cautious. The next few days matter: if the open weights arrive on 15/10/2026 with a permissive licence, a model card and independent benchmarks that hold up, Step 5 becomes a serious default for cost-sensitive agent work. Until then it is a good thing to test, not yet a thing to bet on. We will update this article when the weights land.
Sources
- eesel AI: StepFun Step 5 Preview, specs, pricing, benchmarks (21/09/2026): StepFun's benchmark table, pricing page figures, platform specs and the charts reproduced above.
- Data Studios: StepFun launches Step 5 Preview: architecture, context, Artificial Analysis index and pricing.
- MarkTechPost: StepFun Launches Step 5 Preview: independent launch summary (not retrieved in full).
- Julian Goldie SEO on YouTube, Hermes Agent + Step 5 and WorldofAI: creator coverage embedded above.
Charts: StepFun's launch graphics as reproduced in eesel AI's review. Both are vendor-produced and not independent measurements.
Last updated: 10/10/2026. Sourced from StepFun launch material as reported by third parties. We have not tested Step 5 ourselves; all benchmark figures are vendor-reported, and open weights were not yet available at the time of writing.
Get the free guide: Claude vs ChatGPT, Gemini & Grok
A 20-page playbook covering everything you need to choose and use the big four AI models in 2026, full cost and feature comparisons, what each is best (and worst) at, and how-tos for images, vectors, building a website, Claude Code and more.






