AI Tools Review
Meta Superintelligence Labs' official abstract hero graphic for the Muse Spark 1.3 announcement, showing intertwined blue and white flowing lines against a light background.

Insights

Meta Muse Spark 1.3 Review: Benchmarks & Pricing

AI Tools Review Editorial Team6 September 2026Updated 6 September 2026
  • Meta
  • Muse Spark
  • Muse Code
  • Meta Superintelligence Labs

Quick answer:

Meta released Muse Spark 1.3 on 2 September 2026, a coding- and agent-focused update to Muse Spark 1.2 with a one-million-token combined context window. Meta's own benchmark chart shows a genuinely split result rather than a clean win: Muse Spark 1.3 (max) leads Claude Opus 5 and GPT-5.6 Sol on coding and long-context evaluations, 75.4% on DeepSWE v1.1 and a tie with GPT-5.6 Sol at 88.8% on Terminal-Bench 2.1, but trails Opus 5 on every broader agent-workflow benchmark Meta published, including GDPVal-AA v2 and JobBench. Independent testing from Artificial Analysis found a similarly mixed picture, with some evaluations improving and others declining, plus roughly 3x the token usage of 1.2 on comparable tasks. Pricing is unchanged: $1.25/$4.25 per million input/output tokens on the standard tier, or a steeply discounted ~$0.10/~$0.20 contributor tier that trains on your data. There are still no open weights.

A month after Muse Spark 1.2 and Muse Code marked Meta's most serious entry yet into agentic coding, still landing a clear second behind Claude Opus 5 on every chart the company published, Meta has come back with a third iteration that changes the shape of the gap rather than closing it outright. Muse Spark 1.3 does not chase Opus 5 across the board; it overtakes it on coding and long-context work specifically, while still trailing on the broader category of real-world agentic tasks.

This review works from Meta's own published benchmark chart on research.meta.ai, independent testing published by Artificial Analysis, and launch-day reporting that scrutinised the pricing and safety claims. Where Meta's numbers and independent numbers disagree, both are shown.

Note: benchmark figures in the chart below are Meta's own published numbers from research.meta.ai; AI Tools Review has not independently reproduced them. The independent-testing section draws on Artificial Analysis's separate evaluation. Prices are billed in US dollars.

A hands-on look at Muse Spark 1.3's benchmarks, agent skills and a live demo building a 3D board game from scratch.

Executive Summary

  • Released 2 September 2026 by Meta Superintelligence Labs, rolling out through Muse Code and the Meta Model API the same day.
  • Coding and long-context wins: 75.4% on DeepSWE v1.1 (ahead of Opus 5's 74.0% and GPT-5.6 Sol's 73.0%), a tie with GPT-5.6 Sol at 88.8% on Terminal-Bench 2.1, and a near-total lead on the two MRCR long-context bands (98.5% and 98.1%).
  • A consistent agent-benchmark gap: Claude Opus 5 (max) leads Muse Spark 1.3 on GDPVal-AA v2, JobBench, DeepSearchQA, the Agentic IF Index and AutomationBench, the five benchmarks Meta groups under "Agent" alongside OSWorld 2.0.
  • Mixed independent results. Artificial Analysis found real gains on some evaluations (Tau3-Bench Banking, Terminal-Bench 2.1) and real declines on others (a long-document test, a 6,000-question knowledge test).
  • Roughly 3x more verbose than Muse Spark 1.2 on comparable tasks, per community and Artificial Analysis testing, which raised the effective cost per completed task despite flat per-token pricing.
  • Unchanged pricing: $1.25/$4.25 standard, ~$0.10/~$0.20 contributor tier that trades a steep discount for training-data rights.
  • Still fully proprietary, with the more capable max reasoning setting gated to partners pending additional safety testing.

Lineage: From Muse Spark 1.2 to 1.3

Muse Spark 1.2 landed on 5 August 2026 alongside Muse Code, Meta's first purpose-built terminal coding agent, and closed part of the gap to the closed frontier while still landing a clear second behind Claude Opus 5 on every benchmark chart Meta published: 82.9% versus 86.7% on Terminal-Bench 2.1, 59.3% versus 65.0% on DeepSWE 1.1. That release was notable for what it lacked as much as what it had: no open weights, no safety card, and a pricing structure that offered a steep discount in exchange for training-data rights.

Muse Spark 1.3 arrives four weeks later with a narrower, more specific set of claims. Rather than a general capability refresh, Meta frames it around sustaining longer agentic workflows and coding tasks specifically, reporting roughly 20% fewer tool calls and 25% fewer tokens used on comparable long-horizon jobs against 1.2, plus a "cleaner overall coding style" with reduced verbosity in code output. The pace of iteration, a third named release in five months, is itself a signal: Meta is treating Muse Spark as a fast-moving product line rather than an annual flagship cycle, the same posture Anthropic and OpenAI have taken with Claude and GPT point releases.

What's Actually New

Muse Spark 1.3 is a closed-weights reasoning model that reads text, images and video and holds a combined one million tokens of context. It ships in two reasoning settings: xhigh, available immediately on both Muse Code and the Meta Model API, and max, a more compute-intensive setting that Meta has restricted to partners while it completes additional safety testing, a gating decision Meta has not applied to any earlier Muse Spark release.

Meta describes three specific behavioural improvements over 1.2. First, the model is meant to "better sustain longer-horizon work by collaborating with users and juggling multiple workflows", with improved constraint preservation across multi-step tasks, the kind of drift where an agent forgets an instruction given thirty steps earlier that has plagued long-running coding sessions industry-wide. Second, Meta reports improved self-awareness of the model's own capabilities and limitations, which shows up in practice as a willingness to say it does not know something rather than guessing, a behaviour independent hands-on testers have specifically called out. Third, Meta reports better multitasking within a single conversation thread, tracking several parallel lines of work without losing track of any one of them.

On the coding side specifically, Meta describes "cleaner overall coding style" with reduced verbosity, which sits in some tension with what independent testers actually measured, covered below. Safety-wise, Meta reports stronger adversarial robustness and improved resistance to prompt injection, plus better calibration on irreversible actions, asking for confirmation before doing anything that cannot be undone.

Benchmarks: A Split Result

Meta published an eleven-panel benchmark scorecard comparing Muse Spark 1.3 (max) against its own predecessor, Muse Spark 1.2 (xhigh), GPT-5.6 Sol (max) and Claude Opus 5 (max), grouped into three categories: Agent, Long Context and Coding. It is the most direct three-way comparison Meta has published to date, and it is genuinely mixed rather than flattering.

Meta's official benchmark scorecard comparing Muse Spark 1.3 (max) against Muse Spark 1.2 (xhigh), GPT-5.6 Sol (max) and Claude Opus 5 (max) across Agent benchmarks (GDPVal-AA v2, JobBench, OSWorld 2.0, DeepSearchQA, Agentic IF Index, AutomationBench), Long Context benchmarks (MRCR 256K-512K and 512K-1M) and Coding benchmarks (DeepSWE v1.1, SWEAtlas CodeBase QnA, Terminal-Bench 2.1)
Meta's official Muse Spark 1.3 benchmark scorecard. Source: research.meta.ai.
BenchmarkMuse Spark 1.3 (max)Muse Spark 1.2 (xhigh)GPT-5.6 Sol (max)Opus 5 (max)
GDPVal-AA v21754161517101824
JobBench64.961.645.465.7
OSWorld 2.0 (partial)66.947.662.768.3
OSWorld 2.0 (binary)32.017.927.331.4
DeepSearchQA89.485.993.090.4
Agentic IF Index57.846.260.559.1
AutomationBench49.438.246.750.3
MRCR 256K–512K98.566.391.5
MRCR 512K–1M98.155.573.8
DeepSWE v1.175.455.073.074.0
SWEAtlas CodeBase QnA59.446.253.552.7
Terminal-Bench 2.188.882.988.886.6

Read the Coding and Long Context rows first, because they are the clearest story on the chart. On DeepSWE v1.1 Muse Spark 1.3 scores 75.4%, ahead of both Opus 5 (74.0%) and GPT-5.6 Sol (73.0%). On SWEAtlas CodeBase QnA, repository-level code comprehension, it leads more clearly still at 59.4% against 52.7% and 53.5%. On Terminal-Bench 2.1 it ties GPT-5.6 Sol exactly at 88.8%, both ahead of Opus 5's 86.6%. And on long-context retrieval, the two MRCR bands, it is not close: 98.5% and 98.1% against GPT-5.6 Sol's 91.5% and 73.8%, with the gap widening sharply as the context window grows toward the full million tokens. Meta does not report an Opus 5 score on either MRCR band.

Now the other half. Across the six benchmarks Meta groups under "Agent", GDPVal-AA v2, JobBench, OSWorld 2.0, DeepSearchQA, the Agentic IF Index and AutomationBench, Claude Opus 5 (max) posts the highest score on every one except OSWorld 2.0's binary success metric, where Muse Spark 1.3's 32.0% edges Opus 5's 31.4%. The margins are not enormous, typically 1 to 4 points, but they are consistent: on GDPVal-AA v2, an Elo-style measure of real-world professional task quality, Opus 5's 1824 leads Muse Spark 1.3's 1754 by 70 points, a gap similar in scale to the one Opus 5 held over the open-weight K2 Horizon flagship on the same benchmark family. GPT-5.6 Sol, meanwhile, takes outright top marks on DeepSearchQA (93.0%) and the Agentic IF Index (60.5%), beating both Meta and Anthropic's entries on deep web research and instruction-following specifically.

The generational jump over Muse Spark 1.2 is unambiguous everywhere on the chart, which is worth separating from the three-way competitive picture: +13.2 points on OSWorld 2.0 (partial), +32.2 on MRCR 256K–512K, +20.4 on DeepSWE v1.1. Whatever else is true about how 1.3 stacks up against Opus 5 and GPT-5.6 Sol, it is a substantial improvement on its own immediate predecessor across every panel Meta published.

What Independent Testing Found

Artificial Analysis ran its own independent evaluation of the xhigh tier rather than relying solely on Meta's published chart, and the results complicate the picture further rather than simply confirming it. Tau3-Bench Banking, an agentic tool-use evaluation, improved sharply from 35% under 1.2 to 47% under 1.3, and Terminal-Bench 2.1 improved from 80% to 85% in Artificial Analysis's own harness, both consistent with Meta's narrative of a genuine coding and tool-use upgrade.

But two other evaluations moved backward. A long-document comprehension test declined from 83% to 79%, and a 6,000-question general knowledge test fell from 45% to 42%. Neither drop is enormous in isolation, but both cut directly against the idea that 1.3 is a strict upgrade on 1.2, and the long-document result sits awkwardly next to Meta's own MRCR long-context numbers, which show a near-total lead. The likely explanation is that MRCR specifically tests needle-in-haystack retrieval, finding a known fact planted in a long context, while Artificial Analysis's long-document test likely demands broader comprehension across the whole document, a different and harder skill that a model can fail even while acing retrieval.

The most consequential independent finding is verbosity. Community testing and Artificial Analysis both report Muse Spark 1.3 using roughly three times the output tokens of 1.2 on comparable tasks, directly at odds with Meta's own claim of "cleaner overall coding style" with reduced verbosity. That gap between the vendor's framing and independent measurement is large enough to flag rather than gloss over. The practical consequence is cost: Artificial Analysis's composite pricing index for the model rose from around $0.40 to around $0.55 per task despite the per-token rate staying completely flat, because a model that writes three times as much per response costs roughly three times as much per completed job, benchmark scores aside.

A weekly AI news round-up covering Muse Spark 1.3 alongside GPT-6 Astra, Gemini 3.8 Flash and Claude Fable 5.1, released the same week.

Safety Claims and What Backs Them

Meta states that Muse Spark 1.3 offers "more resistance to adversarial inputs and prompt injection" and improved calibration around irreversible actions, meaning the model is meant to pause and confirm before doing anything it cannot undo, deleting a file, sending a payment, force-pushing to a shared branch, rather than proceeding on its own initiative. Behaviourally, this lines up with the "asks clarifying questions, calls for help when stuck" framing that also shows up in independent hands-on demos.

What is missing is a published measurement to back either claim. There is no release-specific adversarial-robustness benchmark, no sample size, and no third-party red-teaming report attached to this launch, a gap independent coverage of the release flagged directly. That puts Muse Spark 1.3 in the same position as Muse Spark 1.2 before it: safety-relevant claims stated in prose, without the kind of model card or dangerous-capability evaluation that Anthropic, OpenAI and Google DeepMind now routinely publish alongside comparable frontier releases. The gating of the more capable max tier to partners "pending additional safety testing" is itself the most concrete safety signal in the announcement, since it implies Meta's own internal evaluation found something in the max configuration that needed more scrutiny before broader release, though Meta has not said what.

Pricing: Unchanged, Two Tiers

TierInput ($/1M)Cached input ($/1M)Output ($/1M)Trains on your data?
Standard$1.25$0.15$4.25No
Contributor~$0.10~$0.02~$0.20Yes

Meta carried pricing over unchanged from Muse Spark 1.2: $1.25 per million input tokens and $4.25 per million output tokens on the standard tier, with prompts and completions kept private. The contributor tier remains roughly $0.10/$0.20, a 10 to 20 times discount in exchange for letting Meta train future models on your traffic, still with the same tighter rate limits that make it a fit for individual experimentation rather than production workloads.

Because the price per token has not moved, the verbosity increase documented above is the real story on cost. A team billed on the standard tier that saw its median response length triple between 1.2 and 1.3 will see its bill move by roughly the same factor for the same workload, independent of anything in the benchmark chart. Anyone migrating from 1.2 to 1.3 in production should budget for that before assuming "same price, better model" is the whole story.

Limitations

  • Trails Opus 5 on agent-workflow benchmarks. Claude Opus 5 (max) leads on GDPVal-AA v2, JobBench, OSWorld 2.0 (partial), DeepSearchQA and the Agentic IF Index; GPT-5.6 Sol tops DeepSearchQA and the Agentic IF Index outright.
  • Roughly 3x more verbose than 1.2 on comparable tasks per independent testing, directly contradicting Meta's "cleaner, reduced verbosity" framing and raising real-world cost per task despite flat per-token pricing.
  • Mixed independent results, not a clean upgrade. Artificial Analysis found declines on a long-document test and a general knowledge test alongside gains on Tau3-Bench Banking and Terminal-Bench 2.1.
  • No published safety benchmark. Adversarial-robustness and prompt-injection-resistance claims are stated in prose with no disclosed methodology, sample size or third-party red-teaming.
  • Max reasoning tier gated to partners pending unspecified additional safety testing, with no public timeline for broader release.
  • No open weights or licence, continuing the fully proprietary posture Meta adopted with the original Muse Spark in April 2026.

How Muse Spark 1.3 Compares

Against Claude Opus 5, Muse Spark 1.3 has flipped the script from 1.2: it now wins on coding-specific and long-context evaluations (DeepSWE v1.1, SWEAtlas CodeBase QnA, Terminal-Bench 2.1, both MRCR bands) while still losing on the broader category of real-world agentic and professional-workflow tasks (GDPVal-AA v2, JobBench, DeepSearchQA, the Agentic IF Index, AutomationBench). Which model "wins" now genuinely depends on the workload: a team building a coding agent that works inside a huge repository has good reason to prefer 1.3's numbers, while a team automating broader professional workflows still has good reason to prefer Opus 5's.

Against GPT-6 Astra, a much larger and more expensive frontier release from OpenAI announced the same week, the comparison is closer to apples-to-oranges: Astra targets frontier reasoning and long-horizon autonomy at a correspondingly higher price point, while Muse Spark 1.3 is explicitly positioned as a cheaper, coding-and-agent-focused option. Against GPT-5.6 Sol specifically, the model Meta chose to chart directly, Sol takes outright top marks on DeepSearchQA and the Agentic IF Index but ties Muse Spark 1.3 on Terminal-Bench 2.1 and trails it clearly on both coding-comprehension and long-context evaluations.

Against Qwen 3.8 Max and the open-weight K2 Horizon fleet, Muse Spark 1.3's advantage is entirely about being a named, closed-API product from a well-resourced lab with day-one availability, rather than a licensing or self-hosting story; neither of those releases offers Muse Spark 1.3's combination of price and coding-benchmark position, but both offer something it still does not: a downloadable model.

Who Should Use It

Worth trying now if your workload is genuinely coding- or repository-comprehension-heavy: the DeepSWE, SWEAtlas and Terminal-Bench numbers are a real, Meta-published lead over both Opus 5 and GPT-5.6 Sol, and the long-context numbers suggest it will hold up well on large codebases specifically. Budget-conscious teams comfortable with the contributor tier's training-data trade-off for prototyping get a genuinely capable model at a fraction of frontier pricing.

Better to wait or test carefully if your workload looks more like broad professional-task automation than coding specifically, where Opus 5's lead on GDPVal-AA v2, JobBench and the Agentic IF Index is the more relevant signal. And before migrating a production coding workload straight from 1.2 to 1.3, run your own cost comparison rather than assuming flat pricing means a flat bill: the roughly 3x verbosity increase independent testers found means your actual spend per completed task may rise even as the per-token rate does not move.

The Bottom Line

Muse Spark 1.3 is a genuine, unambiguous upgrade over Muse Spark 1.2 on every benchmark Meta published, and a specifically targeted, credible challenger to Claude Opus 5 and GPT-5.6 Sol on coding and long-context work rather than a broad frontier contender. That is a narrower and more honest claim than most model launches make, and Meta's own chart, showing Opus 5 still ahead on every agent-workflow benchmark, backs it up rather than contradicting it.

The open question independent testing raises is whether the coding win is as clean in practice as the benchmark chart suggests, given the roughly threefold increase in output verbosity Artificial Analysis and community testers both measured. A model that writes three times as much to reach a slightly higher DeepSWE score has not necessarily become three times more useful, and the real-world cost-per-completed-task comparison matters as much as the headline percentage for any team actually deciding what to switch to.

Meta's official Muse Spark 1.3 announcement carries the full benchmark scorecard referenced throughout this article.

Last updated: 6 September 2026, four days after Muse Spark 1.3 launched. This article will be revised if Meta publishes a safety card, expands access to the max reasoning tier, or if further independent evaluations land.

Free Guide

Get the free guide: Claude vs ChatGPT, Gemini & Grok

A 20-page playbook covering everything you need to choose and use the big four AI models in 2026, full cost and feature comparisons, what each is best (and worst) at, and how-tos for images, vectors, building a website, Claude Code and more.

Pop your email in to get it free
Preview of the free guide: Claude vs ChatGPT, Gemini and Grok, 2026 features, pricing and what-you-can-do comparison.

Frequently Asked Questions

What is Muse Spark 1.3?
Muse Spark 1.3 is Meta Superintelligence Labs' update to its Muse Spark line of closed-weights reasoning models, released 2 September 2026. It reads text, images and video, holds a combined one million tokens of context, and is aimed at longer agentic workflows and coding tasks than its predecessor, Muse Spark 1.2. It ships in two reasoning tiers: xhigh, available immediately through Muse Code and the Meta Model API, and max, a more compute-intensive setting limited to partners while Meta completes additional safety testing.
How does Muse Spark 1.3 perform against Claude Opus 5 and GPT-5.6 Sol?
It is a split result, and the split is the real story. On Meta's own published chart, Muse Spark 1.3 (max) leads on coding and long-context evaluations: 75.4% on DeepSWE v1.1 versus 74.0% for Claude Opus 5 (max) and 73.0% for GPT-5.6 Sol (max), 59.4% on SWEAtlas CodeBase QnA versus 52.7% and 53.5%, and 98.5%/98.1% on the two MRCR long-context bands versus 66.3%/55.5% for its own predecessor. On Terminal-Bench 2.1 it ties GPT-5.6 Sol at 88.8%, ahead of Opus 5's 86.6%. But on the broader agent-workflow benchmarks, GDPVal-AA v2, JobBench, DeepSearchQA, the Agentic IF Index and AutomationBench, Claude Opus 5 (max) leads on every one, typically by 1 to 4 points.
How much does Muse Spark 1.3 cost?
Pricing is unchanged from Muse Spark 1.2. The standard tier costs $1.25 per million input tokens, $0.15 per million cached input tokens and $4.25 per million output tokens, with a commitment that prompts and completions are not used for training. A separate contributor tier costs roughly $0.10 input and $0.20 output per million tokens, 10 to 20 times cheaper, in exchange for letting Meta train on your traffic.
Is Muse Spark 1.3 open source?
No. Like Muse Spark 1.2 before it, Muse Spark 1.3 is entirely proprietary: no downloadable weights, no licence. Meta has not commented further on open-sourcing since Mark Zuckerberg's non-committal 'I'll have more to share on that soon' when asked about Muse Code in August 2026.
What do independent tests say about Muse Spark 1.3?
Artificial Analysis's independent testing of the xhigh tier found a mixed picture rather than a clean generational win: Tau3-Bench Banking improved from 35% to 47% and Terminal-Bench 2.1 from 80% to 85%, but a long-document test declined from 83% to 79% and a 6,000-question knowledge test fell from 45% to 42%. Community testing also reports Muse Spark 1.3 using roughly three times the output tokens of 1.2 for comparable tasks, which pushed Artificial Analysis's composite cost index up from around $0.40 to around $0.55 despite the per-token price staying flat.
AI Tools Review Editorial Team

AI Tools Review Editorial Team Expert verified

Our editorial team consists of veteran AI researchers, software engineers, and industry analysts. We spend hundreds of hours benchmarking frontier models natively to provide you with objective, actionable intelligence on agentic AI capabilities and cybersecurity landscapes.