Muse Spark 1.2 Benchmarks
Meta has not published benchmark scores at launch. Here is what to look for on the vendor page — and how to run your own eval.
Muse Spark 1.2 benchmark scores were not published at Meta's 2026-08-05 launch. There are no MMLU, SWE-bench, HumanEval, MATH, GPQA, or LiveCodeBench numbers to cite.
This page names the evals that matter for a coding-grade model, tells you what to look for on the vendor page, and shows you how to run a workload eval that beats any generic benchmark for your team.
Anything below that claims a specific score would be fabricated. This page is deliberately eval-shaped, not score-shaped.
What Meta Has Not Published
Meta has not published benchmark scores for Muse Spark 1.2 at launch. Verify at https://developer.meta.com/ai/products/muse-code/.
There are no numbers to cite for MMLU (general reasoning), SWE-bench (real-world coding), HumanEval (code generation), MATH (math reasoning), GPQA (graduate-level QA), or LiveCodeBench (competitive programming).
There is also no head-to-head chart against Muse Spark 1.1, Claude Fable 5, or GPT-5.6. Any published elsewhere without a Meta primary source should be treated as unverified.
- MMLU: not published
- SWE-bench: not published
- HumanEval: not published
- MATH / GPQA / LiveCodeBench: not published
- Head-to-head charts: not published
Waiting on Muse Spark 1.2 benchmarks to make a call? Book a consult and we will run a workload eval that beats any public score.
Book a ConsultationWhat to Look For on the Vendor Page
When Meta does publish evals, check SWE-bench Verified first — it is the closest proxy for real-world coding-agent value.
For multi-agent coordination claims, look for eval categories that measure long-horizon task success (e.g., end-to-end refactor completion, multi-file bug fixes) rather than single-turn HumanEval accuracy.
For the Contributor tier claim of parity with Standard, watch for a same-model-different-tier comparison. Meta may or may not publish this, but it is the honest way to prove the discount does not degrade output.
- SWE-bench Verified: real-world coding proxy
- Long-horizon task success > single-turn HumanEval
- Same-model-different-tier comparison for the Contributor claim
- Reject unofficial scores without a Meta source
Run Your Own Workload Eval
A workload eval beats a public benchmark for team decisions. Pick 20 real tasks from your last two sprints (bug fixes, feature slices, PR reviews, refactors) and run each through Muse Spark 1.2 alongside your current agent.
Score on four axes: correctness (does the change pass CI), completeness (did the agent finish or ask for help), edit locality (did it touch only what it should), and time-to-green (elapsed time to a passing PR).
When we run our own competitor-monitor routine we score new tools on this exact rubric because SWE-bench averages hide the failure modes teams actually care about. Twenty tasks is enough to see the pattern.
- 20 real tasks from recent sprints
- Score: correctness, completeness, edit locality, time-to-green
- Compare against your current agent, not against a paper number
- Twenty tasks is usually enough to see a pattern
What the Launch Claims Imply
Meta positions Muse Spark 1.2 as "optimized for long-horizon, multi-agentic workflows and transparent auditability". Read that as an implicit claim that SWE-bench-style single-turn HumanEval is not the right measurement for this model.
If Meta publishes a benchmark suite, expect it to lead with multi-turn agent completion rates and end-to-end task success. That is the story Meta needs to tell to justify a new model behind a new agent.
Until those scores exist, workload evals are the only defensible input for a buying decision. Treat marketing claims as hypotheses, not conclusions.
- Meta implies long-horizon > single-turn
- Expect multi-turn completion rates when scores land
- Workload evals are the only defensible input today
- Marketing claims are hypotheses, not conclusions
How to use Muse Spark 1.2
You do not host Muse Spark 1.2 yourself — you use it through a tool, so "getting started" really means choosing the right one.
The fastest way to put Muse Spark 1.2 to work day to day is inside an AI IDE, and Cursor is the most popular — it supports it directly, so you can be working in minutes. Prefer a different editor? Windsurf, Zed, and GitHub Copilot drive these models too.
Frequently Asked Questions
- Meta has not published benchmark scores at launch. Verify at https://developer.meta.com/ai/products/muse-code/.
- SWE-bench scores are not published. When they land, SWE-bench Verified is the number to compare against.
- There are no head-to-head benchmark scores at launch. Run a workload eval on your own tasks to decide.
- Meta has not published a same-model-different-tier comparison. The tiers use the same model, but a formal parity eval is not documented.
- The vendor page at https://developer.meta.com/ai/products/muse-code/ is the primary source. Reject unofficial scores without a Meta reference.
- Twenty real tasks from your last two sprints is usually enough to see a pattern in correctness, completeness, edit locality, and time-to-green.
Need a Workload Eval on Muse Spark 1.2?
We design 20-task workload evals that give teams honest go/no-go data in weeks. Book a free 30-minute audit and we will scope one for your codebase.
Book a Free Audit