AI Business Benchmarks: Which AI Actually Does the Job
The benchmark built for businesses, not researchers. We hand real jobs to every major AI and score one thing: could you send the result as-is?
Marketing
4 jobs| Job | Use | Done | Cost | Burned you | DoubleCheck It? |
|---|---|---|---|---|---|
| Make my writing not sound like AIA stiff draft that reads like a robot wrote it | GrokGrok 4.5 · xAI | 82/100 | ~$0.06 | 3/100 | No |
Who is being tested
Cost per try is estimated from each vendor's published token price for a job this size. | |||||
| Write my blog postA topic, an outline and your sources | ChatGPTGPT-5.6 Sol · OpenAI | 97/100 | ~$0.05 | 0/100 | Skim |
Who is being tested
Cost per try is estimated from each vendor's published token price for a job this size. | |||||
| Write my email newsletterThis month's news and links | ClaudeFable 5.1 · Anthropic | 88/100 | ~$0.21 | 2/100 | Skim |
Who is being tested
Cost per try is estimated from each vendor's published token price for a job this size. | |||||
| Turn one post into social postsOne blog post and the platforms you post on | ClaudeFable 5.1 · Anthropic | 87/100 | ~$0.18 | 3/100 | No |
Who is being tested
Cost per try is estimated from each vendor's published token price for a job this size. | |||||
Ads & creative
4 jobs| Job | Use | Done | Cost | Burned you | DoubleCheck It? |
|---|---|---|---|---|---|
| Make an ad videoA script and a photo of the product | DeepSeekV4 Pro | 85/100 | ~$0.04 | 3/100 | No |
Who is being tested
| |||||
| Put my product in a new photoA real photo of your product | ElevenLabsv3 · voice | 97/100 | ~$0.10 | 0/100 | No |
Who is being tested
| |||||
| Record a voiceoverA script and the voice you want | ClaudeFable 5.1 · Anthropic | 95/100 | ~$0.32 | 1/100 | No |
Who is being tested
| |||||
| Make a talking-head videoA script and a presenter | ElevenLabsv3 · voice | 93/100 | ~$0.21 | 2/100 | No |
Who is being tested
| |||||
Sales
4 jobs| Job | Use | Done | Cost | Burned you | DoubleCheck It? |
|---|---|---|---|---|---|
| Write a cold email sequenceYour offer and the people you are writing to | ClaudeFable 5.1 · Anthropic | 95/100 | ~$0.25 | 1/100 | No |
Who is being tested
Cost per try is estimated from each vendor's published token price for a job this size. | |||||
| Look up a prospect before a callA company name | Gemini3.6 Flash · Google | 94/100 | ~$0.02 | 1/100 | Skim |
Who is being tested
Cost per try is estimated from each vendor's published token price for a job this size. | |||||
| Write a proposalThe client's brief and your pricing | ClaudeFable 5.1 · Anthropic | 80/100 | ~$0.16 | 3/100 | No |
Who is being tested
Cost per try is estimated from each vendor's published token price for a job this size. | |||||
| Follow up on a quoteA quote that has gone quiet for ten days | Gemini3.6 Flash · Google | 94/100 | ~$0.03 | 1/100 | Skim |
Who is being tested
Cost per try is estimated from each vendor's published token price for a job this size. | |||||
Customer service
4 jobs| Job | Use | Done | Cost | Burned you | DoubleCheck It? |
|---|---|---|---|---|---|
| Answer a customer emailThe email and your help docs | ChatGPTGPT-5.6 Sol · OpenAI | 97/100 | ~$0.11 | 0/100 | Skim |
Who is being tested
Cost per try is estimated from each vendor's published token price for a job this size. | |||||
| Decide whether to give a refundThe complaint and your refund policy | ElevenLabsv3 · voice | 97/100 | ~$0.10 | 1/100 | Skim |
Who is being tested
Cost per try is estimated from each vendor's published token price for a job this size. | |||||
| Write a help articleFive tickets you've already answered | ClaudeFable 5.1 · Anthropic | 95/100 | ~$0.17 | 1/100 | Skim |
Who is being tested
Cost per try is estimated from each vendor's published token price for a job this size. | |||||
| Answer the phoneYour script and a caller who interrupts | ClaudeFable 5.1 · Anthropic | 97/100 | ~$0.12 | 1/100 | Skim |
Who is being tested
Cost per try is estimated from each vendor's published token price for a job this size. | |||||
Bookkeeping & admin
4 jobs| Job | Use | Done | Cost | Burned you | DoubleCheck It? |
|---|---|---|---|---|---|
| Fix a broken spreadsheetA workbook with broken formulas | ChatGPTGPT-5.6 Sol · OpenAI | 96/100 | ~$0.09 | 1/100 | Skim |
Who is being tested
Cost per try is estimated from each vendor's published token price for a job this size. | |||||
| Type up a pile of invoicesTwenty messy supplier invoices | ChatGPTGPT-5.6 Sol · OpenAI | 97/100 | ~$0.12 | 0/100 | No |
Who is being tested
Cost per try is estimated from each vendor's published token price for a job this size. | |||||
| Match the books to the bankYour books and the bank export | Gemini3.6 Flash · Google | 83/100 | ~$0.02 | 2/100 | Skim |
Who is being tested
Cost per try is estimated from each vendor's published token price for a job this size. | |||||
| Build the monthly reportRaw numbers and last month's report | Gemini3.6 Flash · Google | 92/100 | ~$0.06 | 1/100 | Skim |
Who is being tested
Cost per try is estimated from each vendor's published token price for a job this size. | |||||
Hiring & staff
4 jobs| Job | Use | Done | Cost | Burned you | DoubleCheck It? |
|---|---|---|---|---|---|
| Write a job adThe role and the pay range | ClaudeFable 5.1 · Anthropic | 85/100 | ~$0.21 | 3/100 | No |
Who is being tested
Cost per try is estimated from each vendor's published token price for a job this size. | |||||
| Sort through a pile of resumesForty resumes and your criteria | Gemini3.6 Flash · Google | 93/100 | ~$0.04 | 1/100 | Skim |
Who is being tested
Cost per try is estimated from each vendor's published token price for a job this size. | |||||
| Write a performance reviewA year of your notes | ChatGPTGPT-5.6 Sol · OpenAI | 91/100 | ~$0.09 | 2/100 | Skim |
Who is being tested
Cost per try is estimated from each vendor's published token price for a job this size. | |||||
| Plan a new hire's first monthThe role and the tools you use | ChatGPTGPT-5.6 Sol · OpenAI | 97/100 | ~$0.03 | 0/100 | No |
Who is being tested
Cost per try is estimated from each vendor's published token price for a job this size. | |||||
Scores updated Sep 30, 2026 · 24 jobs scored · 18 AIs tested · prices from the Layer3Labs model price ledger
Law firm
5 jobs| Job | Tested on | Cost per try |
|---|---|---|
| Review a contract against our playbookThe contract and your negotiation playbook | 12 AIs + Harvey | $0.01 – $0.58 |
Who is being tested
Cost per try is estimated from each vendor's published token price for a job this size. | ||
| Abstract a leaseA commercial lease | 12 AIs + Harvey | $0.01 – $0.58 |
Who is being tested
Cost per try is estimated from each vendor's published token price for a job this size. | ||
| Find the red flags in a data roomA data room | 12 AIs + Harvey | $0.01 – $0.58 |
Who is being tested
Cost per try is estimated from each vendor's published token price for a job this size. | ||
| Run a title chainThe deeds and recorded documents | 12 AIs + Harvey | $0.01 – $0.58 |
Who is being tested
Cost per try is estimated from each vendor's published token price for a job this size. | ||
| Redline a contractThe contract and your position | 12 AIs + Harvey | $0.01 – $0.58 |
Who is being tested
Cost per try is estimated from each vendor's published token price for a job this size. | ||
Accounting & tax
5 jobs| Job | Tested on | Cost per try |
|---|---|---|
| Reconcile two ledgersTwo ledgers that do not agree | 12 AIs | $0.01 – $0.58 |
Who is being tested
Cost per try is estimated from each vendor's published token price for a job this size. | ||
| Tie out the financial statementsThe statements and the trial balance | 12 AIs | $0.01 – $0.58 |
Who is being tested
Cost per try is estimated from each vendor's published token price for a job this size. | ||
| Research a tax positionThe facts and the question | 12 AIs | <$0.01 – $0.37 |
Who is being tested
Cost per try is estimated from each vendor's published token price for a job this size. | ||
| Run payrollHours, pay rates and the state | 12 AIs | <$0.01 – $0.16 |
Who is being tested
Cost per try is estimated from each vendor's published token price for a job this size. | ||
| Find unusual expensesA year of expenses and your threshold | 12 AIs | $0.01 – $0.58 |
Who is being tested
Cost per try is estimated from each vendor's published token price for a job this size. | ||
Consulting
5 jobs| Job | Tested on | Cost per try |
|---|---|---|
| Size a marketThe market and your sources | 12 AIs + Hebbia | <$0.01 – $0.37 |
Who is being tested
Cost per try is estimated from each vendor's published token price for a job this size. | ||
| Write up a competitorA competitor name | 12 AIs + Hebbia | <$0.01 – $0.37 |
Who is being tested
Cost per try is estimated from each vendor's published token price for a job this size. | ||
| Recommend a priceYour costs, market and goals | 12 AIs + Hebbia | <$0.01 – $0.16 |
Who is being tested
Cost per try is estimated from each vendor's published token price for a job this size. | ||
| Write the board reportRaw numbers and your format | 12 AIs + Hebbia | <$0.01 – $0.37 |
Who is being tested
Cost per try is estimated from each vendor's published token price for a job this size. | ||
| Build the pitch storyThe data room | 12 AIs + Hebbia | $0.01 – $0.58 |
Who is being tested
Cost per try is estimated from each vendor's published token price for a job this size. | ||
Compliance & risk
5 jobs| Job | Tested on | Cost per try |
|---|---|---|
| Run a gap assessmentYour controls and the framework: GDPR, HIPAA or SOC 2 | 12 AIs + Vanta | $0.01 – $0.58 |
Who is being tested
Cost per try is estimated from each vendor's published token price for a job this size. | ||
| Write a policy to a frameworkThe framework and how you work | 12 AIs + Vanta | <$0.01 – $0.37 |
Who is being tested
Cost per try is estimated from each vendor's published token price for a job this size. | ||
| Work out what a rule change means for usThe rule change and your obligations | 12 AIs + Vanta | <$0.01 – $0.37 |
Who is being tested
Cost per try is estimated from each vendor's published token price for a job this size. | ||
| Screen a customer for AML/KYCCustomer details and your screening rules | 12 AIs + Vanta | <$0.01 – $0.16 |
Who is being tested
Cost per try is estimated from each vendor's published token price for a job this size. | ||
| Assemble the ESG reportThe underlying data | 12 AIs + Vanta | $0.01 – $0.58 |
Who is being tested
Cost per try is estimated from each vendor's published token price for a job this size. | ||
Clinic (paperwork only)
5 jobs| Job | Tested on | Cost per try |
|---|---|---|
| Write the visit noteThe visit transcript | 12 AIs + Abridge | <$0.01 – $0.16 |
Who is being tested
Cost per try is estimated from each vendor's published token price for a job this size. | ||
| Code the visitThe visit note | 12 AIs + Abridge | <$0.01 – $0.16 |
Who is being tested
Cost per try is estimated from each vendor's published token price for a job this size. | ||
| Put together a prior-auth packetThe chart and the payer's criteria | 12 AIs + Abridge | $0.01 – $0.58 |
Who is being tested
Cost per try is estimated from each vendor's published token price for a job this size. | ||
| Check a credentialing fileA provider's credentialing file | 12 AIs + Abridge | $0.01 – $0.58 |
Who is being tested
Cost per try is estimated from each vendor's published token price for a job this size. | ||
| Write a denial appealThe denial and the chart | 12 AIs + Abridge | <$0.01 – $0.37 |
Who is being tested
Cost per try is estimated from each vendor's published token price for a job this size. | ||
Insurance agency
4 jobs| Job | Tested on | Cost per try |
|---|---|---|
| Assess a submissionThe submission and your guidelines | 12 AIs | $0.01 – $0.58 |
Who is being tested
Cost per try is estimated from each vendor's published token price for a job this size. | ||
| Summarize a claim fileThe claim file | 12 AIs | $0.01 – $0.58 |
Who is being tested
Cost per try is estimated from each vendor's published token price for a job this size. | ||
| Compare two policiesTwo policy documents | 12 AIs | $0.01 – $0.58 |
Who is being tested
Cost per try is estimated from each vendor's published token price for a job this size. | ||
| Draft a coverage letterThe policy and the claim | 12 AIs | <$0.01 – $0.37 |
Who is being tested
Cost per try is estimated from each vendor's published token price for a job this size. | ||
Construction
5 jobs| Job | Tested on | Cost per try |
|---|---|---|
| Take off quantities from drawingsThe drawings | 12 AIs | $0.01 – $0.58 |
Who is being tested
Cost per try is estimated from each vendor's published token price for a job this size. | ||
| Check a submittal against the specThe submittal and the spec | 12 AIs | $0.01 – $0.58 |
Who is being tested
Cost per try is estimated from each vendor's published token price for a job this size. | ||
| Put together a permit packageThe project and the jurisdiction | 12 AIs | <$0.01 – $0.37 |
Who is being tested
Cost per try is estimated from each vendor's published token price for a job this size. | ||
| Write a change orderThe change and your rates | 12 AIs | <$0.01 – $0.16 |
Who is being tested
Cost per try is estimated from each vendor's published token price for a job this size. | ||
| Check plans against the safety codeThe plans and the code | 12 AIs | $0.01 – $0.58 |
Who is being tested
Cost per try is estimated from each vendor's published token price for a job this size. | ||
Bids & procurement
5 jobs| Job | Tested on | Cost per try |
|---|---|---|
| Answer an RFPThe RFP and your capability library | 12 AIs + Hebbia | $0.01 – $0.58 |
Who is being tested
Cost per try is estimated from each vendor's published token price for a job this size. | ||
| Build a compliance matrixThe RFP requirements | 12 AIs + Hebbia | $0.01 – $0.58 |
Who is being tested
Cost per try is estimated from each vendor's published token price for a job this size. | ||
| Score the vendorsVendor responses and your weighting | 12 AIs + Hebbia | $0.01 – $0.58 |
Who is being tested
Cost per try is estimated from each vendor's published token price for a job this size. | ||
| Review a procurement contractThe contract and your policy | 12 AIs + Hebbia | $0.01 – $0.58 |
Who is being tested
Cost per try is estimated from each vendor's published token price for a job this size. | ||
| Decide whether to bidThe opportunity and your criteria | 12 AIs + Hebbia | <$0.01 – $0.37 |
Who is being tested
Cost per try is estimated from each vendor's published token price for a job this size. | ||
The Layer3Labs AI Business Benchmarks test 12 AI models and 10 specialist tools on 63 business jobs: 24 everyday jobs every business has, and 39 specialist jobs from law firms, accounting practices, clinics, insurance agencies, construction firms and compliance teams. No existing AI benchmark measures this work.
How Every Job Is Graded
Every result lands on one of four grades, whether the job is a social post or a prior-authorization packet. It is the same test a business owner applies to work that comes back to their desk.
| Grade | What it means |
|---|---|
| Done | Send it. No edits. |
| Fix | Usable after five or ten minutes of edits. |
| Redo | Wrong. Do it yourself. |
| Burned you | Wrong in a way you would not have caught before it went out. |
What We Measure
- Done rate: how many times out of 100 the work comes back needing no edits.
- Burned-you rate: how often the work comes back wrong in a way you would never catch: an invented figure, a policy you don't have, a case that doesn't exist. Every other benchmark scores what AI gets right. This one also scores what it gets wrong without telling you.
- Cost per finished job: what one usable piece of work costs at published prices, counting the attempts you would throw away.
- DoubleCheck It? Yes, skim or no. Each job runs five times in a row; if the AI doesn't slip and doesn't burn you, you can stop checking.
How We Test
Every job comes with the inputs a real business would hand over: the messy invoices, the refund policy, the lease, the data room. Each AI gets the same inputs and the same instructions, five times. Jobs with a single right answer, like a spreadsheet total or a refund decision under a written policy, are graded mechanically. Every result is checked against the facts in its inputs, so anything invented is caught.
Prices come from each vendor's published rates, tracked on our AI model pricing page. When a price changes there, the cost on every job here changes with it.
Frequently Asked Questions
- The AI Business Benchmarks score AI models and tools on 63 real business jobs, such as writing a newsletter, deciding a refund, typing up invoices or abstracting a lease. Each job is scored on one question: could you send the result without fixing it? The benchmarks are published by Layer3Labs.
- Those benchmarks measure coding puzzles, exam questions and which chat answer people prefer. None of them measure whether AI can do ordinary business work to a standard you would send to a customer. The AI Business Benchmarks test the jobs a business actually hands off, graded the way a business owner judges work.
- The Done rate is how many times out of 100 the work comes back needing no edits at all. Every job on this page states exactly what Done means for that job, so the standard is the same for every AI tested.
- The Burned-you rate is how often the work comes back wrong in a way you would not catch before it went out: an invented figure, a policy your business does not have, a case or citation that does not exist. It is the number that decides whether you can hand a job off without checking every result.
- Cost comes from each vendor's published token prices, tracked on our AI model pricing page, applied to the size of each job. Cost per finished piece of work also counts the attempts you would throw away, so a cheap model that fails often does not look cheaper than it is.
- Every job is run five times on every AI. That shows whether a result is repeatable, which is what decides whether you can stop checking the work, not whether it got lucky once.
- Yes. AI labs can submit a model for scoring ahead of launch and receive an official score they can cite on release day. Email partners@layer3labs.io with the model name and launch window.
- No. The clinic jobs cover paperwork only: visit notes, coding, prior authorizations, credentialing and appeals. No job asks an AI to diagnose, triage or recommend treatment.
Want to Know Which Jobs AI Can Take Off Your Plate?
We map the jobs in your business that AI can do to a send-it-as-is standard, and the ones it can't yet. Start with a free AI workflow audit.
Book a Consultation