Independent reference

Muse Spark 1.2: Full Specs and Benchmarks Behind Muse Code

Muse Spark 1.2 is the frontier model Meta Superintelligence Labs released on 5 August 2026 alongside Muse Code, its terminal coding agent. This page is a sourced spec sheet and benchmark table: what Muse Spark 1.2 actually publishes, what it does not, and where independent testing disagrees with Meta's own numbers.

Model
Muse Spark 1.2, from Meta Superintelligence Labs
Release date
5 August 2026 (1.1 shipped 9 July 2026)
Context window
1M tokens
Input types
Text, image, video, PDF
Output speed
~165 tokens/sec (Artificial Analysis)
Parameters
Not published
Weights
Not released
Paired agent
Muse Code

Muse Spark 1.2 Spec Sheet

Every published Muse Spark 1.2 specification, blanks included, for the model that Muse Code runs on.

SpecificationValueSource
DeveloperMeta Superintelligence LabsMeta AI Research launch post
Release date5 August 2026Official blog
PredecessorMuse Spark 1.1, released 9 July 2026Official blog
Context window1M tokensMeta + Artificial Analysis
Multimodal inputImage, video and PDF alongside textMeta
Output speed~165 tokens/secArtificial Analysis
Parameter countNot published
Knowledge cutoffNot published
Open weightsNot released
Paired harnessMuse Code, co-trained with the modelMeta
AccessMeta Model API, OpenRouter, or Muse CodeMeta + OpenRouter listing

Two fields people search for most — muse spark parameters and knowledge cutoff — genuinely have no published value. Anyone quoting a number for either is guessing. Meta's own model listing sits at developer.meta.com.

Muse Spark Benchmarks, Regime 1: Artificial Analysis (Independent)

Independent Muse Spark benchmark results for the model inside Muse Code, measured by Artificial Analysis.

BenchmarkMuse Spark 1.2Muse Spark 1.1Direction
Intelligence Index5451Improved
Terminal-Bench v2.180%78%Improved
GDPval (Elo)+260 vs 1.1baselineImproved
SciCodeRegressed (no figure published)Worse
Humanity's Last ExamRegressed (no figure published)Worse
Output throughput~165 tokens/sec

All figures in this table come from Artificial Analysis, an independent evaluator. Do not average them with the Meta-reported table below — different harnesses, different prompts, different numbers.

Muse Spark Benchmarks, Regime 2: Meta-Reported (Second-Hand)

Meta's own claimed Muse Spark scores behind Muse Code, relayed by press rather than a primary results page.

BenchmarkMuse Spark 1.2 (Meta-reported)Named rivalRival score
Terminal-Bench 2.182.9%Claude Opus 586.7%
Terminal-Bench 2.182.9%GPT-5.6 Terra81.8%
DeepSWE59.3Claude Opus 565.0

These are as reported by press coverage of the Muse Code launch, not independently reproduced, and are kept separate from the Artificial Analysis table on purpose. Meta's comparison also names GPT-5.6 Terra rather than the flagship Sol — a framing choice commenters on Hacker News picked up on immediately.

Muse Spark 1.1 vs 1.2: What Improved and What Got Worse

A muse spark 1.1 vs 1.2 comparison for Muse Code users that reports regressions as plainly as gains.

One month separates the two releases, and the delta is real but narrow. Independently, Muse Spark 1.2 lifts the Intelligence Index from 51 to 54, Terminal-Bench v2.1 from 78% to 80%, and gains roughly 260 Elo on GDPval. For the agentic coding work Muse Code puts on it, the Terminal-Bench movement matters most.

The unflattering half rarely gets reported: Artificial Analysis also recorded regressions on SciCode and Humanity's Last Exam. Neither comes with a published figure, so all anyone can honestly say is that Muse Spark 1.2 is worse than 1.1 on both. A model tuned hard for terminal agent work giving ground on scientific coding and broad knowledge is an unsurprising trade — but still a trade.

Is Muse Spark Open Source? Weights, Hugging Face and Licence

Muse Spark open weights, Hugging Face and licensing for Muse Code users — the short answer is none exist.

Muse Spark 1.2 is not open source and not open weights. No checkpoint has been published, there is no Muse Spark repository on Hugging Face, and no licence governs local use because there is nothing to run locally — Muse Spark is a hosted service only.

The question stays alive because of Meta's history with Llama. Asked directly about releasing weights, Mark Zuckerberg said he would have more to share on that soon, as reported by The Register — an acknowledgement that the question is open, not a commitment to a date. Until a checkpoint appears, treat every "Muse Spark open weights" claim as speculation.

The same applies to the Muse Code client, which ships as a prebuilt binary with no published source. Neither half of the stack is inspectable today.

How to Access Muse Spark 1.2

Three routes reach Muse Spark 1.2 today, and only one of them involves installing Muse Code at all.

First, the Meta Model API serves Muse Spark 1.2 directly, with per-tier rate limits and per-token billing. Second, OpenRouter lists the model as meta/muse-spark-1.2, the usual route into third-party harnesses. Third, installing Muse Code gets you the model wrapped in Meta's own agent.

Keys, sign-up, rate limits by tier and playground access are covered on our Muse Spark API guide. Token prices for both tiers live on the Muse Code pricing page, with a cost calculator for your own volume.

Before you budget: Muse Spark 1.2 has no Bedrock listing and no free tier, so every route starts with a payment method.

How Muse Spark 1.2 Powers Muse Code

Muse Spark 1.2 and Muse Code were trained together, which shapes what the benchmark numbers above mean.

Meta's launch post states that Muse Spark 1.2 was co-trained with the Muse Code harness rather than adapted to it afterwards. That is why its strongest published result is Terminal-Bench — a benchmark about driving a terminal — and why weaker results sit on tasks Muse Code never asks of it.

Three Muse Spark properties do visible work inside the agent. The 1M-token context lets Muse Code carry a large repository plus a long action history without constant re-reading. Multimodal input makes the video-to-website demo possible. And roughly 165 tokens per second is why long agent runs finish in hours rather than days — a speed advantage commenters on Hacker News rated around three times faster than typical DeepSeek providers.

What Muse Spark does not fix is anything above it in the stack: platform coverage, sign-in, spending caps and skills belong to the agent. Those live on our Muse Code explainer; the pairing against rivals is covered on Muse Code vs Claude Code and Muse Code vs Codex.

Where each thread from this Muse Spark 1.2 spec sheet continues in more depth.

Frequently Asked Questions

What is the Muse Spark context window?

1M tokens. Muse Spark 1.2 accepts up to a million tokens of context, and that window carries multimodal input too — images, video and PDFs, not just text. It is what allows Muse Code to keep a large repository and a long action history in view at once.

How many parameters does Muse Spark have?

Meta has not published a parameter count. Neither the launch post nor the model listing states a size, and no credible third-party estimate exists; the same gap applies to the knowledge cutoff. Any specific parameter figure quoted for Muse Spark 1.2 is invented — treat the field as unknown until Meta publishes a model card.

Is Muse Spark open source?

No. Weights are not released, there is no Muse Spark Hugging Face repository, and no local-use licence exists. Access runs through Meta's hosted API, resellers or Muse Code only. Zuckerberg said he would have more to share on open weights soon, as reported by The Register, but no checkpoint and no date have followed.

What is the Muse Spark release date?

Muse Spark 1.2 launched on 5 August 2026, the same day as Muse Code. Its predecessor, Muse Spark 1.1, shipped on 9 July 2026 — under a month earlier. Two frontier releases in five weeks suggests the version you benchmark today may not be current for long.

Is Muse Spark good for coding?

On terminal agent work, yes — with caveats. Terminal-Bench v2.1 puts Muse Spark at 80% independently and 82.9% by Meta's own reported figure, competitive but below Claude Opus 5 either way. It regressed on SciCode versus 1.1, so scientific code generation is its weaker flank. For long runs inside Muse Code it is a strong pick.

Where does Muse Spark appear on a leaderboard?

Artificial Analysis is the main independent tracker, scoring Muse Spark 1.2 at 54 on its Intelligence Index versus 51 for 1.1. OpenRouter also publishes usage-based rankings for the meta/muse-spark-1.2 listing. Meta's own Terminal-Bench and DeepSWE claims for Muse Code are relayed through press coverage, so keep those separate from any leaderboard number.

Is Muse Spark 1.2 actually better than 1.1?

Better for agents, mixed elsewhere. Muse Spark 1.2 gains on the Intelligence Index, Terminal-Bench and GDPval, but Artificial Analysis also logged regressions on SciCode and Humanity's Last Exam without figures. If your workload is long Muse Code sessions, upgrade; if it is scientific coding or broad knowledge, test before assuming a win.

Does Muse Code use any model other than Muse Spark 1.2?

No. Muse Code ships bound to Muse Spark 1.2 with no published model-selection option, unlike agents that let you swap in third-party models. To run Muse Spark in a different harness the traffic flows the other way — through OpenRouter into tools like opencode or Cline, per our API guide.

Sources

  1. Meta AI Research — Introducing Muse Code and Muse Spark 1.2research.meta.ai
  2. Artificial Analysis — Muse Spark 1.2 evaluationartificialanalysis.ai
  3. Meta developer — Muse Spark model listingdeveloper.meta.com
  4. OpenRouter — meta/muse-spark-1.2openrouter.ai
  5. Hacker News — Muse Code and Muse Spark 1.2 discussionnews.ycombinator.com