Reviewed by Jonathan West · Updated Jul 17, 2026

Muse Spark 1.2 Benchmarks

Meta has not published benchmark scores at launch. Here is what to look for on the vendor page — and how to run your own eval.

Reviewed by Jonathan West · Updated Jul 17, 2026

Muse Spark 1.2 benchmark scores were not published at Meta's 2026-08-05 launch. There are no MMLU, SWE-bench, HumanEval, MATH, GPQA, or LiveCodeBench numbers to cite.

This page names the evals that matter for a coding-grade model, tells you what to look for on the vendor page, and shows you how to run a workload eval that beats any generic benchmark for your team.

Anything below that claims a specific score would be fabricated. This page is deliberately eval-shaped, not score-shaped.


What Meta Has Not Published

Meta has not published benchmark scores for Muse Spark 1.2 at launch. Verify at https://developer.meta.com/ai/products/muse-code/.

There are no numbers to cite for MMLU (general reasoning), SWE-bench (real-world coding), HumanEval (code generation), MATH (math reasoning), GPQA (graduate-level QA), or LiveCodeBench (competitive programming).

There is also no head-to-head chart against Muse Spark 1.1, Claude Fable 5, or GPT-5.6. Any published elsewhere without a Meta primary source should be treated as unverified.

  • MMLU: not published
  • SWE-bench: not published
  • HumanEval: not published
  • MATH / GPQA / LiveCodeBench: not published
  • Head-to-head charts: not published

Waiting on Muse Spark 1.2 benchmarks to make a call? Book a consult and we will run a workload eval that beats any public score.

Book a Consultation

What to Look For on the Vendor Page

When Meta does publish evals, check SWE-bench Verified first — it is the closest proxy for real-world coding-agent value.

For multi-agent coordination claims, look for eval categories that measure long-horizon task success (e.g., end-to-end refactor completion, multi-file bug fixes) rather than single-turn HumanEval accuracy.

For the Contributor tier claim of parity with Standard, watch for a same-model-different-tier comparison. Meta may or may not publish this, but it is the honest way to prove the discount does not degrade output.

  • SWE-bench Verified: real-world coding proxy
  • Long-horizon task success > single-turn HumanEval
  • Same-model-different-tier comparison for the Contributor claim
  • Reject unofficial scores without a Meta source

Run Your Own Workload Eval

A workload eval beats a public benchmark for team decisions. Pick 20 real tasks from your last two sprints (bug fixes, feature slices, PR reviews, refactors) and run each through Muse Spark 1.2 alongside your current agent.

Score on four axes: correctness (does the change pass CI), completeness (did the agent finish or ask for help), edit locality (did it touch only what it should), and time-to-green (elapsed time to a passing PR).

When we run our own competitor-monitor routine we score new tools on this exact rubric because SWE-bench averages hide the failure modes teams actually care about. Twenty tasks is enough to see the pattern.

  • 20 real tasks from recent sprints
  • Score: correctness, completeness, edit locality, time-to-green
  • Compare against your current agent, not against a paper number
  • Twenty tasks is usually enough to see a pattern

What the Launch Claims Imply

Meta positions Muse Spark 1.2 as "optimized for long-horizon, multi-agentic workflows and transparent auditability". Read that as an implicit claim that SWE-bench-style single-turn HumanEval is not the right measurement for this model.

If Meta publishes a benchmark suite, expect it to lead with multi-turn agent completion rates and end-to-end task success. That is the story Meta needs to tell to justify a new model behind a new agent.

Until those scores exist, workload evals are the only defensible input for a buying decision. Treat marketing claims as hypotheses, not conclusions.

  • Meta implies long-horizon > single-turn
  • Expect multi-turn completion rates when scores land
  • Workload evals are the only defensible input today
  • Marketing claims are hypotheses, not conclusions

How to use Muse Spark 1.2

You do not host Muse Spark 1.2 yourself — you use it through a tool, so "getting started" really means choosing the right one.

The fastest way to put Muse Spark 1.2 to work day to day is inside an AI IDE, and Cursor is the most popular — it supports it directly, so you can be working in minutes. Prefer a different editor? Windsurf, Zed, and GitHub Copilot drive these models too.

Frequently Asked Questions

  • Meta has not published benchmark scores at launch. Verify at https://developer.meta.com/ai/products/muse-code/.
  • SWE-bench scores are not published. When they land, SWE-bench Verified is the number to compare against.
  • There are no head-to-head benchmark scores at launch. Run a workload eval on your own tasks to decide.
  • Meta has not published a same-model-different-tier comparison. The tiers use the same model, but a formal parity eval is not documented.
  • The vendor page at https://developer.meta.com/ai/products/muse-code/ is the primary source. Reject unofficial scores without a Meta reference.
  • Twenty real tasks from your last two sprints is usually enough to see a pattern in correctness, completeness, edit locality, and time-to-green.

Need a Workload Eval on Muse Spark 1.2?

We design 20-task workload evals that give teams honest go/no-go data in weeks. Book a free 30-minute audit and we will scope one for your codebase.

Book a Free Audit
Disclosure: Layer3Labs is reader-supported. When you buy through links on this page we may earn an affiliate commission, at no extra cost to you. Our picks are chosen on the merits — commissions never influence the ranking.