Reviewed by Jonathan West · Updated Sep 5, 2026

Claude Sonnet 5 Benchmarks: What Anthropic Has Published

The qualitative claim Anthropic makes for Sonnet 5, the numbers it has not released, and how rival models score on the same evaluations.

Reviewed by Jonathan West · Updated Sep 5, 2026

Anthropic has not published SWE-bench Verified, CursorBench, or AA (Artificial Analysis) Intelligence Index scores for Claude Sonnet 5. That means there is no verified head-to-head benchmark comparing it with its closest rivals.

We test new models on a client's actual tasks before recommending one. Sonnet 5 is a clear example of why: sometimes, a published leaderboard simply cannot tell you what you need to know.

Anthropic released Sonnet 5 on June 30, 2026, describing its quality as close to the previous Opus tier but at a lower price. However, that remains a qualitative claim, which carries less weight in a procurement decision than a measured benchmark score.


What Anthropic Has Published for Sonnet 5

Anthropic calls Sonnet 5 its most agentic Sonnet model yet. It plans tasks, uses tools such as a browser or terminal, and completes multi-step work with less hand-holding than earlier Sonnet models needed.

That is a positioning claim, not a score. Anthropic states that Sonnet 5's quality sits close to the prior Opus 4.8 tier, at a lower price and on a 1-million-token context window in beta, but none of that is a number you can plug into a comparison table.

One detail is confirmed, though. Sonnet 5 runs on a newer tokenizer that produces roughly 30 percent more tokens for the same text than the prior generation did, even though Anthropic has not published matching performance scores alongside that change.

  • Positioning claim: quality close to the prior Opus tier, at a lower price
  • Context window: 1M tokens in beta on the API
  • Confirmed tokenizer change, without a matching benchmark release

Need a real answer on how Sonnet 5 performs on your workload? We run the test on your own tasks instead of citing a leaderboard.

Book a Consultation

What Anthropic Has Not Published

Three scores are missing. Anthropic has not released a SWE-bench Verified score, a CursorBench score, or an AA Intelligence Index score for Sonnet 5 as of this writing, and those are the three benchmarks most often used to compare current-generation models on coding and agentic reasoning.

That gap shows up directly in cross-model comparisons. A Sonnet 5 vs Grok 4.6 comparison can state Grok's published CursorBench score of 69.9 percent, but it has no Sonnet 5 number to put beside it, only Anthropic's general quality claim.

A missing number is not a bad one. It just means the comparison cannot be made from public data alone. Treat any specific score you see attributed to Sonnet 5 on one of these benchmarks as unconfirmed until Anthropic publishes it directly.

  • No published SWE-bench Verified score for Sonnet 5
  • No published CursorBench score for Sonnet 5
  • No published AA Intelligence Index score for Sonnet 5

What Rival Models Have Published on the Same Benchmarks

xAI's Grok 4.6 publishes real numbers: 61 on the AA Intelligence Index, 69.9 percent on CursorBench v3.2, and 65.9 percent on DeepSWE v1.1, all in its own release materials.

OpenAI's GPT-5.6 Sol does too. It matches Grok 4.6's AA Intelligence Index score of 61 and posts 67.2 percent on CursorBench v3.2, per the same comparison sources.

Zhipu's GLM 5.2 and Moonshot AI's Kimi K3 take a different route entirely. Both ship as open-weight models, so a team can run its own benchmark suite directly against the released weights instead of relying on a vendor's self-reported number.

Sonnet 5 is the outlier here. It is closed and unscored on these particular evaluations, which is a gap in what Anthropic has released, not a claim that Sonnet 5 performs worse.

  • Grok 4.6: 61 AA Intelligence Index, 69.9% CursorBench v3.2, 65.9% DeepSWE v1.1
  • GPT-5.6 Sol: 61 AA Intelligence Index, 67.2% CursorBench v3.2
  • GLM 5.2 and Kimi K3: open weights, so you can run your own benchmark directly
  • Sonnet 5: closed model, no matching published score on any of the three

Why the Missing Number Matters for Procurement

A procurement or compliance process that requires a cited, third-party benchmark score has nothing to cite for Sonnet 5 on SWE-bench Verified, CursorBench, or the AA Intelligence Index.

That does not make Sonnet 5 a bad choice. It changes the argument you bring to a budget review. The justification has to rest on price, default availability, and your own test results instead of a leaderboard entry.

That could shift. Our answer would change the moment Anthropic publishes a matching score on any of these three benchmarks. Until then, treat every specific number you see attributed to Sonnet 5 on them as unconfirmed.

  • No citable third-party benchmark score for Sonnet 5 today
  • Justify a Sonnet 5 choice on price, availability, and your own tests instead
  • Revisit this once Anthropic publishes a matching score

How to Test Sonnet 5 Yourself

Start with your own tasks, not a puzzle. Pull five to ten real jobs from your own workflow, the kind you would actually hand a model day to day.

Run the same tasks on Sonnet 5 and on whichever model you are weighing it against. Score each on accuracy, speed, and cost per task rather than a single pass or fail.

One test beat every leaderboard for us. Running a client's own failing test suite through a candidate model on a coding-heavy build told us more in an afternoon than any published score did, because it surfaced the specific failure patterns that mattered for that codebase.

  • Use 5-10 of your own real tasks, not a generic benchmark puzzle
  • Score accuracy, speed, and cost per task, not just pass/fail
  • Your own test suite surfaces failure patterns a leaderboard cannot

Why a Vendor Withholds a Benchmark Number

A vendor usually skips a benchmark for one of three reasons: the model scores worse than its predecessor on that specific test, the benchmark measures a capability outside the model's intended job, or the number simply was not ready by launch day and gets added later.

Anthropic has not said which applies to Sonnet 5's missing SWE-bench Verified and CursorBench scores. Reading the silence as a hidden weak score is a guess either way, not a fact you can act on.

Watch for the number arriving late. Vendors sometimes publish a missing benchmark score weeks or months after launch, once internal evaluation catches up, so a gap on launch day is not necessarily permanent.

  • A missing score usually means a weak result, a mismatched test, or a late number
  • Anthropic has not stated which reason applies to Sonnet 5
  • A missing score at launch can still arrive later

How to use Claude Sonnet 5

You do not host Claude Sonnet 5 yourself — you use it through a tool, so "getting started" really means choosing the right one.

The fastest way to put Claude Sonnet 5 to work day to day is inside an AI IDE, and Cursor is the most popular — it supports it directly, so you can be working in minutes. The maker's own option is Claude Code for Claude Sonnet 5, if you want the native experience. Prefer a different editor? Windsurf, Zed, and GitHub Copilot drive these models too.

Frequently Asked Questions

  • No. Anthropic has not published a SWE-bench Verified score for Sonnet 5 as of this writing, so no verified comparison to a model that does publish one, such as GPT-5.6 Sol, currently exists.
  • That comparison cannot be made from public data. Grok 4.6 publishes a CursorBench v3.2 score of 69.9%, but Anthropic has not released a matching score for Sonnet 5.
  • Anthropic has not stated a reason. It has published a qualitative claim, that Sonnet 5's quality sits close to the prior Opus tier, without a matching third-party benchmark score.
  • Both ship as open-weight models, which lets you run your own benchmark suite directly against the released weights instead of relying on a self-reported vendor score.
  • Run 5 to 10 of your own real tasks on both models and compare accuracy, speed, and cost per task. A published leaderboard cannot answer a question Anthropic has not scored.
  • Anthropic has not said. If it publishes a matching SWE-bench Verified or CursorBench score, a direct comparison to rivals like Grok 4.6 or GPT-5.6 Sol becomes possible for the first time.

Test Sonnet 5 Against Your Own Tasks

Book a free 30-minute AI workflow audit. At Layer3Labs, we run your real tasks against Sonnet 5 and its closest rivals and hand you a comparison the public benchmarks cannot give you.

Book Your Free Audit
Disclosure: Layer3Labs is reader-supported. When you buy through links on this page we may earn an affiliate commission, at no extra cost to you. Our picks are chosen on the merits — commissions never influence the ranking.