Reviewed by Jonathan West · Updated Sep 25, 2026

Tau Bench Explained

How Sierra's benchmark scores AI agents that must use tools and follow policy with a simulated customer.

Reviewed by Jonathan West · Updated Sep 25, 2026

Tau bench is an evaluation benchmark developed by Sierra that measures how effectively artificial intelligence (AI) agents follow business policies and call software tools while conversing with simulated users.

At Layer3Labs, we audit enterprise AI agent workflows and examine how lab benchmark scores map to daily customer interactions.

A high score on a standard single-turn benchmark rarely guarantees that an agent will update databases correctly or adhere to complex service rules over prolonged conversations.


What τ-bench Measures

The τ-bench framework tests conversational agents by pairing them with a language model user simulator across multi-turn service scenarios.

Each task equips the agent with domain-specific application programming interface (API) tools and an explicit policy manual.

Evaluation inspects whether the agent follows policy guidelines while resolving the simulated customer's request.

The user simulator runs gpt-4-0613 to generate dynamic conversational replies.

Scoring does not rely on subjective text evaluations from another model.

Instead, the benchmark calculates a strict mathematical reward based on database state and output precision.

Success requires that the final database state is identical to the unique ground truth outcome database.

The agent's responses must also communicate all necessary information to the user.

The benchmark defines reward as r = r_action × r_output ∈ {0,1}, meaning a run receives zero if an action fails or required output details are omitted.

Initial testing covered two distinct service environments.

The τ-retail environment contains 115 tasks focused on orders, returns, and address modifications.

The τ-airline domain includes 50 tasks handling flight cancellations, booking changes, and passenger itineraries.

  • User conversations driven dynamically by the gpt-4-0613 simulator
  • Domain-specific API tools for database modifications
  • Annotated policy rules governing customer permissions and constraints
  • Binary scoring requiring exact final database parity and complete output information
Reward requires both an exact final database match and complete conversational output. A single missing detail or incorrect database edit drops the trial score to zero.

The Pass^k Metric and Why Repeat Trials Matter

The pass^k metric measures the probability that an artificial intelligence agent completes every single trial successfully across k independent attempts.

This calculation contrasts sharply with the traditional pass@k metric common in software coding benchmarks.

The pass@k metric scores whether at least one attempt out of k independent trials succeeds.

Customer service requires reliability.

A coding benchmark can count a task as solved if one of several tries works, but a customer sees every try.

In the τ-bench paper, gpt-4o solved under 50% of tasks on one try, and its pass^8 in retail fell below 25%.

Sierra reported that GPT-4o's retail score fell about 60% from pass^1 to pass^8, so a single-try figure overstates how the agent behaves across repeat runs.

In our client engagements, we examine this performance decay when scoping customer service agent architectures.

  • pass^k calculates the probability that all k independent trials succeed, averaged across tasks
  • pass@k calculates the chance that at least one trial out of k succeeds
  • gpt-4o scored under 50% on pass^1 in retail and dropped below 25% on pass^8
  • Sierra documented a 60% decline in GPT-4o reliability when moving from single trials to eight repeat runs
Sierra measured GPT-4o's retail score falling about 60% between pass^1 and pass^8, so ask any AI vendor for both numbers.

Benchmark Versions from Retail to Voice

The benchmark suite has expanded across three major iterations to evaluate broader enterprise environments.

The original τ-bench was submitted to arXiv in June 2024 by Shunyu Yao, Noah Shinn, Pedram Razavi, and Karthik Narasimhan.

It tested agents in retail across 115 tasks and airline customer support across 50 tasks.

The tasks in the original repository are frozen.

A notice on the official sierra-research/tau-bench repository advises researchers to use τ³-bench for updated tasks and new domains.

In June 2025, researchers introduced τ²-bench in a paper submitted to arXiv.

Authors Victor Barres, Honghua Dong, Soham Ray, Xujie Si, and Karthik Narasimhan created a dual-control environment.

In dual control, both the agent and the simulated user use tools in a shared environment modeled as a decentralized partially observable Markov decision process (Dec-POMDP).

The domain models Telecom tasks such as fixing data connections, resolving multimedia messaging service (MMS) issues, and switching network modes.

The authors report that performance drops when agents move from solo execution to dual control.

Sierra's research blog reported a drop of up to 25 points in task success rate when agents operated interactively.

Released in March 2026, τ³-bench introduced Banking and Voice domains on taubench.com.

The τ-Voice paper submitted to arXiv by Soham Ray, Keshav Dhandhania, Victor Barres, and Karthik Narasimhan evaluated full-duplex conversational voice agents.

While GPT-5 reasoning reached 85%, voice agents achieved only 31 to 51% under clean audio conditions and 26 to 38% under realistic acoustic conditions.

The τ-Knowledge paper submitted to arXiv by Quan Shi, Alexandra Zytek, Pedram Razavi, Karthik Narasimhan, and Victor Barres tested agents across roughly 700 interconnected knowledge documents.

Frontier models with high reasoning budgets achieved only ~25.5% pass rates in that banking environment.

  • τ-bench (June 2024): 115 retail tasks and 50 airline tasks evaluated against exact database states
  • τ²-bench (June 2025): dual-control telecom tasks where both agent and simulated user operate tools in a shared environment
  • τ³-bench (March 2026): enterprise banking with roughly 700 knowledge documents alongside full-duplex voice evaluations
  • Voice performance gap: voice agents achieved 31 to 51% under clean audio and 26 to 38% under realistic conditions

Reading the Tau Bench Leaderboard on taubench.com

The official leaderboard on taubench.com reports Pass^1 scores across text, banking knowledge retrieval, and voice environments.

The website displays no evaluation dates.

Any analysis must note the access date directly.

On the τ²-bench Text leaderboard covering Retail, Airline, and Telecom, Qwen3.5-397B-A17B scored 87.9%, Gemini 3.0 Pro scored 85.4%, and Claude Opus 4.5 reached 85.3% as displayed on taubench.com, accessed 2026-09-25.

On the τ³-Banking Text leaderboard for knowledge retrieval, Qwen 3.8 Max achieved 55.2%, Claude Opus 5 reached 48.7%, and Grok 4.5 from xAI scored 47.9% as displayed on taubench.com, accessed 2026-09-25.

On the τ³-Voice leaderboard across all four domains, gpt-live-1 from OpenAI led with 81.7%, followed by Pine Voice Preview at 75.4% and grok-voice-think-fast-1.0 at 67.3% as displayed on taubench.com, accessed 2026-09-25.

Historical scores on the original GitHub repository show different baselines.

On those outdated tasks, Claude 3.5 Sonnet scored 0.460 Pass^1 for airline and 0.692 Pass^1 for retail.

  • taubench.com scores report Pass^1 success rates rather than multi-trial pass^k metrics
  • The public leaderboard displays no run dates, requiring citation dates for verification
  • Top τ²-bench Text scores exceed 85%, while no model shown tops 56% on τ³-Banking
  • Legacy GitHub repository numbers represent obsolete task sets superseded by τ³-bench

Limits of the Tau Bench Method

Tau bench checks the final database state exactly, but simulated testing environments cannot replicate all production customer behaviors.

The primary advantage of the method is objective, deterministic verification.

Because scoring inspects database rows and required text outputs, it eliminates the subjective scoring bias common in model-based judges.

When an AI vendor reports pass^k, intermittent failures show up in the score.

Simulated users introduce clear limitations.

The user simulator runs gpt-4-0613, which cannot fully reflect real human confusion, hostility, or unexpected conversational pivots.

Furthermore, enterprise tasks remain restricted to retail, airline, telecom, and banking scenarios.

An agent that updates an airline reservation may struggle with custom enterprise resource planning systems.

Public leaderboards also prioritize Pass^1 numbers over repeat reliability.

  • Pro: Deterministic database matching removes subjective evaluation bias
  • Pro: pass^k penalizes intermittent agent failures across repeat trials
  • Pro: Policy manuals force agents to respect real business permissions
  • Con: Simulated language model users do not reproduce genuine customer emotional volatility
  • Con: Public leaderboards emphasize single-trial Pass^1 over repeat Pass^8 consistency
  • Con: Legacy GitHub repository tasks are frozen and superseded by newer suites

Research Origins and Authors at Sierra

The τ-bench suite came out of research at Sierra, which SiliconANGLE describes as an AI startup led by former OpenAI board chair Bret Taylor.

Authors on the original project included Shunyu Yao, Noah Shinn, Pedram Razavi, and Karthik Narasimhan.

Subsequent papers brought contributions from Victor Barres, Honghua Dong, Soham Ray, Xujie Si, Keshav Dhandhania, Quan Shi, and Alexandra Zytek.

Noah Shinn was a research scientist at Sierra, per his Sierra author page.

He co-authored the original benchmark and later founded Instinct, an AI personal assistant.

For more details on his research and career, read our profile on Noah Shinn.

Shinn was first author of Reflexion, presented at NeurIPS 2023.

Reflexion established verbal reinforcement learning, allowing agents to improve through linguistic feedback rather than neural network weight updates.

Reflexion achieved 91% pass@1 accuracy on the HumanEval coding benchmark, surpassing the previous state-of-the-art GPT-4 at 80%.

While τ-bench isolates API tool usage and policy enforcement, our guide to AI agent benchmarks compares other agent benchmarks, including WebArena.


Enterprise Evaluation Criteria for Vendor Claims

Enterprise software buyers should check three things whenever an AI vendor cites a tau bench score.

Ask which benchmark release produced the score.

Scores derived from the frozen 2024 repository do not represent current dual-control or banking complexity.

Check that the tested domain matches your own use case.

On taubench.com (accessed 2026-09-25), the top τ²-bench Text scores are above 85%, while the τ-Knowledge paper reports only ~25.5% for frontier models on its roughly 700-document banking tasks.

Ask for pass^k results as well as the Pass^1 figure.

A single-trial success rate conceals whether an agent will fail across repeated interactions.

Teams deploying simple single-turn informational chatbots that retrieve text without writing to databases do not need tau bench.

Those teams should measure retrieval accuracy with standard retrieval augmented generation benchmarks instead.

Our focus on pass^k over pass^1 would change if an agent operates in an environment where human supervisors approve every database change before execution.

In supervised human-in-the-loop workflows, single-turn accuracy matters more than autonomous repeat reliability.

Audit your conversational agent against multi-turn domain requirements before relying on a vendor tau bench score for procurement decisions.

  • Check that the score comes from current tasks on taubench.com rather than the frozen 2024 repository
  • Match the evaluated domain to your specific operational environment rather than accepting generic retail scores
  • Require pass^8 or pass^k metrics to verify repeat consistency before deploying agents to production databases
  • Assess whether human-in-the-loop supervision mitigates multi-turn decay for critical database updates

Frequently Asked Questions

  • A tau bench (τ-bench) is a benchmark developed by Sierra that evaluates how artificial intelligence agents handle multi-turn customer conversations, use software tools, and follow business policies. It scores agents by testing whether the final database state matches an annotated ground truth outcome and whether necessary information was communicated to a simulated user.
  • Tau benches are used to test whether AI agents can reliably execute real-world enterprise tasks without making errors in company databases or violating operational guidelines. Organizations use them to measure tool-calling accuracy, policy adherence, and multi-turn consistency across retail, airline, telecom, and banking environments.
  • A tau 2 bench (τ²-bench) is an updated evaluation framework released in June 2025 that tests conversational agents in dual-control environments where both the agent and the simulated user use tools in a shared environment. Modeled as a decentralized partially observable Markov decision process in a telecom domain, it evaluates collaborative tasks such as fixing data connections, resolving multimedia messaging service issues, and switching network modes.
  • The tau bench leaderboard is a public performance ranking hosted at taubench.com that displays Pass^1 success rates for frontier language models across text, banking knowledge retrieval, and voice benchmarks. As displayed on taubench.com, accessed 2026-09-25, top text models on τ²-bench reach above 85%, while complex tasks like τ³-Banking knowledge retrieval drop frontier models below 56%.
  • The primary pros of tau bench are deterministic database verification and the pass^k repeat consistency metric, while the main cons are simulated user limitations and narrow domain coverage. It removes subjective scoring bias and measures reliability across repeat trials, but its gpt-4-0613 user simulator cannot fully capture human emotional unpredictability, and legacy tasks in the initial repository are no longer updated.

Evaluate Your AI Agent Architecture

Book a 30-minute consultation with Layer3Labs. We review your agent workflows, assess benchmark reliability, and build a production testing framework for your business.

Book a Consultation