All News
metamuse-sparkreleaseagentic-aiopen-weightspricing

Muse Spark 1.3's best results come from a 'max' mode developers can't deploy yet

Meta released Muse Spark 1.3 with frontier-cluster scores, but its best results come from a max mode still in safety testing — and per-task costs rose anyway.

Vlad MakarovVlad Makarovreviewed and published
6 min read
Mentioned models
Muse Spark 1.3's best results come from a 'max' mode developers can't deploy yet

Meta released Muse Spark 1.3 on September 2 — its fourth Muse Spark release in five months, per Artificial Analysis — rolling out across Muse Code and the Meta Model API. Zuckerberg billed it as "the biggest jump we've made so far on coding and agentic work," with "frontier performance almost too cheap to meter." The launch deserves scrutiny beyond the usual roundup: Meta's best results come from a max reasoning configuration developers cannot broadly use yet.

What ships today is Muse Spark 1.3 with the reasoning modes Meta has offered since 1.2, including xhigh; max arrives "shortly," after additional safety testing, per Meta's blog. Artificial Analysis ran max in a limited partner preview and lists no API provider for it. VentureBeat's Carl Franzen pressed the point: not whether Muse Spark 1.3 can touch the frontier, but how close the version companies can deploy today gets — and at what real cost.

The scorecard, split in two

Meta's evaluation report discloses both configurations, so this is not a hidden flagship — the launch materials simply showcase the max variant, where the largest scores live. All figures below are Meta's own, from the launch post and evaluation report:

Benchmarkmaxxhigh (shipping)
GDPval-AA v2 (Elo)17541709
OSWorld 2.066.957.2
JobBench64.961.2
Terminal-Bench 2.188.889.2
DeepSearchQA89.489.4

The split is uneven: max wins the agentic rows by 45 Elo to nearly ten points, while the shipping config takes Terminal-Bench 2.1 and ties DeepSearchQA.

Artificial Analysis scores the pair at 62 (max) and 61 (xhigh) on its Intelligence Index. Shipping xhigh ties GPT-5.6 Sol max, Grok 4.6 high and Claude Opus 5 high — inside the frontier cluster, but not its leader. Anthropic's Claude Fable 5.1, out two days with its own per-task cost controversy, holds 66 at max and 65 at xhigh; Opus 5 sits at 63. Meta closed most of last month's gap — and still matches prior-generation top scores more than it beats current ones.

"Almost too cheap to meter" meets the meter

Zuckerberg's line does not rest on a price cut: standard rates are unchanged from 1.2 — $1.25 per million input, $4.25 output, $0.15 cached. "Too cheap" is a claim about what the tokens buy, not what they cost.

Independent measurement complicates it. Artificial Analysis clocks shipping xhigh at 235.2 output tokens per second and estimates $0.55 per Intelligence Index task at its score of 61, the lowest per-task cost at that intelligence level. Its predecessor ran $0.40 per task at 57; per-task cost rose despite frozen prices, which AA mainly attributes to heavier input-token consumption on agentic evaluations.

That does not contradict Meta's efficiency claims so much as scope them: the roughly 20% fewer tool calls and 25% fewer tokens come from Meta engineers' internal coding comparisons, while AA measures a broader suite of reasoning and agentic tasks. Both can be true — which is why "cheap" gets slippery once models operate as agents.

Meta also kept its Contributor tier — $0.10 in and $0.20 out per million for letting Meta train on prompts and completions, roughly 90% below standard rates. It is the source of the "matching Fable 5 at ten cents in, twenty cents out" figure circulating online.

The same-day context sharpens the economics: Google shipped Gemini 3.8 Flash on September 2, and AA measures it at 59 and $0.58 per task at high reasoning — behind Muse Spark on both, but at roughly 305 output tokens per second versus 235 and an introductory $0.75/$3.75 that expires December 31. Meta chief AI officer Alexandr Wang replied to AA's numbers with "i really hate to say it, but… gemini who?" The measured answer: Gemini is faster today; Muse Spark is the slightly stronger high-effort agent.

Beyond the benchmark rows

Meta's release notes read like a fix-list for agentic operating costs: the model juggles multiple workflows in one long thread, gathers context, asks clarifying questions on ambiguous prompts, confirms before consequential actions, and matches its updates to user preference. Meta also claims better calibration about its limits and stronger prompt-injection resistance. All self-reported — whether enterprise telemetry on retries and interventions confirms any of it at production scale is the open question.

Open weights: "coming soon" is doing heavy lifting

Zuckerberg promises "Muse Spark open weights releases coming soon" and teases a larger model codenamed Watermelon; the blog's roadmap says only "the Muse Spark open weights release" — no version, no date, no license. The Next Web reports Meta has not decided whether to publish 1.3's weights at all, though it still plans to release 1.2's. It compounds an earlier promise: on August 10, the day it open-sourced the 30-billion-parameter Muse Glimmer under Apache 2.0, Meta said Spark 1.2's weights would follow "in the coming weeks." They have not appeared.

In Europe the weights decision is a compliance decision. EU AI Act Article 53 exempts genuinely open general-purpose models from part of its transparency duties — a real license, weights, architecture and usage information public, no non-commercial clause or user thresholds — but the exemption stops at systemic-risk models, which owe the full obligations regardless. Copyright policy and the public training-content summary apply either way; Meta declined to sign the AI Act code of practice in July 2025.

The local crowd noticed the drift: the top r/LocalLLaMA thread on the release — roughly 750 points and 190-plus comments — is a screenshot of Zuckerberg's post, its author waiting for something else: "I am still waiting for Llama 5, because Muse Spark will be too big for me, or just something between Glimmer and Spark."

The "surpassed" chatter, checked

r/singularity ran with "Meta's muse spark 1.3 surpassed fable 5 and GPT 5.6 sol" — about 450 points and 130-plus comments. The attached screenshot — Artificial Analysis' Intelligence Index chart — undercuts the title: Muse Spark 1.3 max scores 62, tied with Claude Fable 5, ahead of GPT-5.6 Sol max at 61. A tie, not a surpass, and that score belongs to the partner-preview configuration; shipping xhigh merely ties GPT-5.6 Sol. The r/ClaudeAI thread claiming Fable-5 parity "with .10 cents input, .20 cents output" refers to that Contributor tier — a discount paid in training rights.

The bottom line

Muse Spark 1.3 is Meta's strongest price-performance release: shipping xhigh sits inside the frontier cluster at roughly a tenth of Anthropic's flagship list prices, and the behavior changes target agents' real cost centers. Unresolved: whether max's extra points survive general availability; what "open weights release" means while 1.2's weights are missing; and whether per-task cost holds now that AA has measured one increase. Meta's efficiency numbers describe its own coding workflows — the meter that matters runs on yours. Full specs and scores are on the Muse Spark 1.3 model page.

Related Articles

Scroll down

to load the next article