AI Tools Review
StepFun Step 5 Preview: Specs, Pricing & Verdict

Insights

StepFun Step 5 Preview: Specs, Pricing & Verdict

AI Tools Review Editorial Team10 October 2026

    Quick Answer:

    Step 5 Preview is StepFun's new flagship, launched on 20/09/2026: a sparse Mixture-of-Experts model with 600 billion total and 27 billion active parameters, a 1 million-token context window and API pricing of $1.00 input and $2.70 output per million tokens. StepFun says it matches the leading open models at about 65% lower cost, and has promised open weights on 15/10/2026. Today it is API-only, every benchmark is self-reported and there is no technical report, so treat it as promising but unproven.

    Most model launches claim to be the best. StepFun's claim is more modest and, for that reason, more interesting: Step 5 Preview aims to be the best model you can get at a given price, not the best model full stop. If the numbers hold, that is exactly the trade-off agent builders care about.

    This guide pulls together StepFun's own launch material as reported by trade and independent sites, plus the Artificial Analysis index figures those reports cite, and the creator coverage that has pushed Step 5 into the Hermes Agent community. We flag every figure as vendor-reported, because that is what it is. We have not run Step 5 ourselves.

    Julian Goldie introduces Step 5 and pitches it as a free Chinese model for agent builders.

    Executive Summary

    What it is: Step 5 Preview (model ID step-5-preview) is a sparse Mixture-of-Experts reasoning model from StepFun, a Shanghai lab that has shipped the "Step" family across text, vision and audio for a couple of years. It is described by StepFun as a frontier model for complex work and agentic tasks, with a lean towards software engineering and finance.

    • Specs: 600B total parameters, 27B active per token (about 4.5%), 1M-token context, up to 64k output tokens, text output only, image input (one review also reports short video clips).
    • Price: $1.00 input (cache miss), $0.05 input (cache hit), $2.70 output per million tokens, as listed on StepFun's pricing page and echoed by Artificial Analysis.
    • Independent signal: Artificial Analysis assigns an Intelligence Index of 44, which one review describes as roughly level with GLM-5.3 and Kimi K3.
    • Open weights: promised for 15/10/2026. As of 10/10/2026 there is no model card, licence or weights.
    • Self-reported benchmarks: DeepSWE v1.1 67.7, StepCodeBench 49.0, ProgramBench 80.5, Terminal-Bench v4 33.3, FrontierFinance 66.4.
    • Caveats: no technical report, StepCodeBench is StepFun's own benchmark, early community reactions flag verbosity, and the 27B active count is heavier than some rival "flash" tiers.

    Our view: Step 5 Preview is a credible bid for the "best value open tier" position, and the cache-hit price is genuinely attractive for agent loops. It is a preview with unaudited numbers, so the sensible move is to test it on your own tasks and wait for the weights and independent evaluations before committing.

    Who StepFun Is and Why Step 5 Matters

    StepFun (Chinese name 阶跃星辰) is one of the Chinese labs that has combined closed API products with a healthy open-weights footprint on Hugging Face. Its smaller Flash line shipped under Apache 2.0, which is why the market is hopeful, though not certain, about the Step 5 licence. Step 5 is the lab's attempt to move from capable mid-tier models to a contender in the broader frontier conversation that already includes GLM-5.3, Kimi K3, Qwen 3.8 Max and the DeepSeek V4 family.

    It also arrives in the same fortnight as Reflection AI's Beam, a 501B open-weight model from the United States. Together they show where the open-model race is heading: very large sparse models with small active footprints, reinforcement learning at scale, and an argument about cost per solved task rather than parameter count.

    Architecture: 600B Total, 27B Active

    In a dense model every parameter runs for every token. In a sparse Mixture-of-Experts model a router chooses a small subset of expert sub-networks for each token, so only a slice of the network fires. Step 5 holds 600 billion parameters but activates 27 billion per token, which is 4.5% of the total and means about 95.5% sit idle for any one token. That is the economic logic behind the pricing: StepFun can charge like a mid-tier model while drawing on a much larger pool of specialised capacity.

    SpecificationStep 5 Preview
    Model IDstep-5-preview
    ArchitectureSparse Mixture-of-Experts
    Total / active parameters600B / 27B
    Context window1M tokens
    Max output64k tokens
    InputText and images (one review also lists short video clips)
    OutputText
    Reasoning controlreasoning_effort: low, medium or high
    API availabilityFrom 20/09/2026
    Open weightsPromised for 15/10/2026

    The 27B figure needs careful reading. It does not make Step 5 behave like a dense 27B model, because the larger parameter pool gives the router access to many more specialised experts. It also does not make deployment equivalent to hosting a 27B model: the full 600B weight set still has to be stored and distributed across hardware, so memory, interconnect bandwidth and routing efficiency stay important once the weights ship. A sceptical Hacker News commenter quoted in one review pointed out that 27B active is heavier than some rival "flash" models, so the efficiency win depends on intelligence per active parameter rather than the active count itself. Independent inference benchmarks will settle that, not launch charts.

    Capabilities Deep Dive

    The 1M-token context

    StepFun positions the million-token window around workloads that accumulate state over time: repositories, specifications, issue histories, logs, test output and previous attempts for software agents; filings and reports for financial research. A representative agent loop runs from objective to context inspection, plan, tool call, result, state update, next action and verification, and context grows with every cycle. A bigger window gives the loop room before anything has to be summarised or moved into an external memory system.

    Long context is not a free lunch. Filling a window increases prefill work and can hurt latency and cost when an application repeatedly sends hundreds of thousands of tokens, and a long window does not guarantee the agent finds relevant information or recovers from failed steps. For production use you still need retrieval, caching and context selection, which is where the 95% cache-hit discount below becomes relevant.

    Reasoning effort and modalities

    Step 5 is a reasoning model with a reasoning_effort control set to low, medium or high. Because output tokens are billed inclusive of the reasoning trace, effort is a direct cost dial: higher effort means a longer internal trace and a bigger bill. On the multimodal side, one review says you can pass up to 60 images per request; output is text only. Reports differ on whether video is supported (one lists text, images and video, others list text and images), so check the platform documentation before designing around it.

    Coding and finance focus

    StepFun leans on software engineering, saying Step 5 can go from software engineering and visual applications to programmable hardware, and that in its expert evaluations about 70% of participants judged it capable of autonomously solving coding tasks of moderately high complexity. That is a survey-style claim from the vendor, not a reproducible benchmark. There is also a Claude Code integration through StepFun's "Step Plan" that exposes the full 1M context, according to the same review, which is a smart way to meet developers where they already work. If you are choosing tooling, our guide to the best AI coding agents compares the harnesses.

    Finance is the other stated focus, and it is where the table below shows Step 5's most distinctive result.

    Benchmarks: What StepFun Reports

    The Pareto-frontier claim

    StepFun's central argument is a chart. It plots Artificial Analysis Intelligence Index against cost per task on a log scale, and draws a new Pareto frontier that Step 5 Preview sits on. The vendor's claim is a 65% lower cost than GLM-5.3 and Kimi K3 at a comparable level of intelligence.

    Chart plotting Artificial Analysis Intelligence Index against cost per task, with Step 5 Preview on a new Pareto frontier above DeepSeek V4 Pro and Minimax-M3 and at about 65 percent lower cost than GLM-5.3 and Kimi K3
    StepFun's own cost-versus-intelligence chart: Step 5 Preview on a "new Pareto frontier", with an annotated 65% cost reduction against GLM-5.3 and Kimi K3, and GPT-6 Astra and Claude Opus 5 at the high-compute end. Vendor-produced chart, not an independent result. Source: StepFun, via eesel AI's review.

    Read it as the pitch, not the verdict. The vertical axis is an independent index, which gives the chart some credibility, but the framing and the cost axis are StepFun's, and the cost depends on how many tokens each model spends. The independent figure is the Intelligence Index of 44 and the $1.00 and $2.70 prices that Artificial Analysis lists. One review describes 44 as roughly matching GLM-5.3 and Kimi K3.

    The comparison table

    BenchmarkStep 5 PreviewKimi K3GLM-5.3GPT-6 AstraClaude Opus 5
    DeepSWE v1.167.767.566.974.174.0
    StepCodeBench49.043.940.261.063.9
    ProgramBench80.577.872.085.482.3
    Terminal-Bench v433.312.641.957.952.3
    FrontierFinance66.462.664.155.069.7

    All numbers above come from StepFun's product page as reproduced by a review. The picture is consistent: at parity with the leading open models, behind the top closed models on raw coding quality, and cheaper. Finance is the standout, where Step 5 edges both open rivals and beats GPT-6 Astra (66.4 against 55.0), though it still trails Claude Opus 5 (69.7). The one clear loss to an open rival is Terminal-Bench v4, where GLM-5.3 scores 41.9 against Step 5's 33.3, though Step 5 is well ahead of Kimi K3's 12.6 on that row. For context on the closed models, see our GPT-6 Astra review and Claude Opus 5 coverage.

    StepCodeBench radar chart across six task categories showing Step 5 Preview ahead of Kimi K3 and GLM-5.3, with an overall average at four attempts of 49.0, 43.9 and 40.2, and Claude Opus 5 at 63.9
    StepCodeBench top-six task categories (feature modification, bug repair, refactoring, documentation generation, performance tuning and code generation): overall average at four attempts is 49.0 for Step 5 Preview, 43.9 for Kimi K3, 40.2 for GLM-5.3 and 63.9 for Opus 5. StepFun's own benchmark and chart, shown as a rank-preserving view. Source: StepFun, via eesel AI's review.

    Why to be cautious

    • Self-reported. There is no technical report or paper yet.
    • Home-field benchmark. StepCodeBench is StepFun's own test, said to cover 553 repositories in 33 languages. Models usually look good on tests their own team designed.
    • Mixed independent signals. An Intelligence Index of 44 is competitive, but it describes a model in the same band as its open rivals, not above them.
    • Community scepticism. Quoted Hacker News reactions include complaints that the demo traces gave away mistakes and that outputs can be padded and verbose. These are anecdotes, not measurements.

    Pricing and Access

    Per 1M tokensInput (cache miss)Input (cache hit)Output
    step-5-preview$1.00 (about £0.75)$0.05 (about £0.04)$2.70 (about £2.03)
    step-3.7-flash (for context)$0.20$0.04$1.15

    The 95% cache-hit discount is the standout. Agentic workloads re-send a large, mostly static system prompt on every step, so cached input at about $0.05 per million tokens changes the economics of long loops. Output at $2.70 covers both reasoning trace and answer, so a chatty high-effort run costs real money. There is a free V0 tier for testing, with 5 concurrent requests and 10 requests per minute, and rate limits scale with top-ups. The price compares well with closed frontier models; for the other cheap options see our reviews of DeepSeek V4.1 Flash and Solar Mini 4. Sterling figures are approximate conversions at roughly £0.75 to the dollar.

    The Open-Weights Countdown

    The timeline matters. StepFun opened paid API access on 20/09/2026 and put a public countdown on its homepage for open weights on 15/10/2026. On launch the Hugging Face repository existed but was an empty placeholder with no weights, model card or licence, according to the reviews we read, and we found no post-release coverage as of 10/10/2026. StepFun's Flash line used Apache 2.0, so a permissive release is likely, but that is an inference. If you are planning on-premises deployment, do not build around weights that have not landed; the date, licence and model card are what to check on 15/10/2026.

    The open-weights promise is also what makes Step 5 comparable to Beam. Both are sparse models of roughly 500 to 600 billion parameters, with active counts in the mid-twenties of billions, both with release dates still ahead of them.

    Using Step 5 with Agents like Hermes

    Step 5 has been picked up quickly by the open agent community. Julian Goldie's video pairing it with Hermes Agent, embedded below, hints at why: an agent harness plus a cheap long-context model is exactly the combination that agent builders want. For background on the harness itself, read our Hermes Agent explainer and the earlier Solar Mini 4 and Hermes piece, which covers another cheap model suited to that setup.

    Julian Goldie pairs Step 5 with Hermes Agent and pitches the free access.

    Two cautions. First, "free" here means the limited V0 tier and promotional access, not unlimited production use. Second, the more context an agent accumulates, the more it costs, so lean on the cache-hit price and keep a retrieval layer. Reasoning-heavy runs at high effort can burn output tokens quickly.

    Real-World Reception Versus Benchmarks

    Early reactions follow a familiar pattern. Supporters like the honest positioning, with one commenter noting that StepFun frames itself as the best of the cheaper tier rather than claiming to top the charts. Critics describe earlier Step models as reasoning too much and too long, and worry about waffly output. Both can be true: a model can win on price per unit of intelligence and still be verbose enough to erode the saving. The WorldofAI news roundup embedded below places Step 5 among the week's launches, and is useful for seeing how it is being framed against the leaks and rumours around it.

    WorldofAI's weekly roundup, including its take on the Step 5 release.

    The practical way to cut through the noise is to run your own evaluation: take 20 to 50 real tasks from your workload, run them at low, medium and high reasoning effort, and record cost, latency and pass rate. That tells you more than any vendor chart, and it is cheap at these prices.

    How It Compares

    • Versus Kimi K3 and GLM-5.3: similar intelligence, lower price, per StepFun and Artificial Analysis. Step 5 loses to GLM-5.3 on Terminal-Bench v4. See Kimi K3 and GLM-5.3.
    • Versus Reflection Beam: Beam is open-weight first and more explicit about its training recipe; Step 5 is API-first and already priced. See our Beam review.
    • Versus GPT-6 Astra and Claude Opus 5: behind on raw coding (DeepSWE 67.7 against 74.1 and 74.0), ahead of Astra on FrontierFinance, and far cheaper.
    • Versus DeepSeek V4.1 Flash: a heavier active count (27B) for a higher-tier claim; see our DeepSeek V4.1 Flash review for the efficiency-first alternative.

    Limitations

    • Preview status. Behaviour, pricing and limits may change before general availability.
    • No technical report, model card or confirmed licence at the time of writing.
    • Self-graded results, including an in-house benchmark.
    • Verbosity and reasoning cost. Output tokens include the trace, so long reasoning raises bills.
    • Heavy to self-host. 600B parameters need serious hardware even with 27B active.
    • Safety evidence. We found no published system card or dangerous-capability evaluation.
    • Data governance. It is a hosted API from a Chinese lab; regulated buyers should check data-handling terms before sending sensitive material.

    Who Should Use It

    Try it if you run agent loops with large, repeated prompts and want the cache-hit price, build coding or finance research tools over big document sets, or want to prepare for an open-weight release you can self-host. Wait if you need a settled model with a model card and independent evaluations, have strict data-residency requirements, or need the strongest possible coding model, where closed frontier models still lead. Test before trusting in every case: the benchmarks are the vendor's.

    The Bottom Line

    Step 5 Preview is a sensible, honestly positioned launch: a large sparse model that tries to match the best open models at a much lower price, with a pricing structure that suits agents. The evidence so far is a self-reported table, an independent index score that puts it in the same band as its rivals, and a set of community reactions that are positive but cautious. The next few days matter: if the open weights arrive on 15/10/2026 with a permissive licence, a model card and independent benchmarks that hold up, Step 5 becomes a serious default for cost-sensitive agent work. Until then it is a good thing to test, not yet a thing to bet on. We will update this article when the weights land.

    Sources

    Charts: StepFun's launch graphics as reproduced in eesel AI's review. Both are vendor-produced and not independent measurements.

    Last updated: 10/10/2026. Sourced from StepFun launch material as reported by third parties. We have not tested Step 5 ourselves; all benchmark figures are vendor-reported, and open weights were not yet available at the time of writing.

    Free Guide

    Get the free guide: Claude vs ChatGPT, Gemini & Grok

    A 20-page playbook covering everything you need to choose and use the big four AI models in 2026, full cost and feature comparisons, what each is best (and worst) at, and how-tos for images, vectors, building a website, Claude Code and more.

    Pop your email in to get it free
    Preview of the free guide: Claude vs ChatGPT, Gemini and Grok, 2026 features, pricing and what-you-can-do comparison.

    Frequently Asked Questions

    What is StepFun Step 5 Preview?
    Step 5 Preview is the new flagship model from Shanghai-based StepFun, launched on 20/09/2026. It is a sparse Mixture-of-Experts model with 600 billion total parameters and 27 billion active per token, a 1 million-token context window, extended reasoning with low, medium and high effort settings, and image input. It is aimed at long-running agentic work, software engineering and finance.
    How much does Step 5 Preview cost?
    StepFun lists $1.00 per million input tokens on a cache miss, $0.05 on a cache hit and $2.70 per million output tokens, with output including the reasoning trace. For comparison, its smaller step-3.7-flash is listed at $0.20 input, $0.04 cached input and $1.15 output. A free tier exists for testing, limited to 5 concurrent requests and 10 requests per minute.
    When will Step 5 be open-weight?
    StepFun has publicly counted down to an open-weights release on 15/10/2026. As of 10/10/2026 the model is API-only; reports say the Hugging Face repository was an empty placeholder with no weights, model card or licence. StepFun's earlier Flash models used Apache 2.0, so a permissive licence is plausible, but it has not been confirmed.
    Is Step 5 better than Kimi K3 and GLM-5.3?
    On StepFun's own table it is level with or slightly ahead of them on most rows: DeepSWE v1.1 at 67.7 against 67.5 for Kimi K3 and 66.9 for GLM-5.3, StepCodeBench at 49.0 against 43.9 and 40.2, and ProgramBench at 80.5 against 77.8 and 72.0. It is behind GLM-5.3 on Terminal-Bench v4 (33.3 against 41.9). All numbers are self-reported and there is no technical report yet.
    Can I use Step 5 with Hermes Agent or Claude Code?
    Creators such as Julian Goldie have shown Step 5 running inside Hermes Agent, and a review of the launch reports a Claude Code integration through StepFun's Step Plan that exposes the full 1M context. Check StepFun's current documentation for set-up details, as the model is a preview and endpoints may change.

    Explore more AI tool comparisons

    In-depth reviews, benchmarks and guides to help you choose the right AI tools.

    Browse all reviews
    AI Tools Review Editorial Team

    AI Tools Review Editorial Team Expert verified

    Our editorial team consists of veteran AI researchers, software engineers, and industry analysts. We spend hundreds of hours benchmarking frontier models natively to provide you with objective, actionable intelligence on agentic AI capabilities and cybersecurity landscapes.