Skip to content
  • Models
  • Rankings
  • Ori
Sign Up
Sign Up
All benchmarks

τ²-Bench: Airline

τ²-Bench Airline tests whether a model can do an airline support agent's job: follow a policy manual, talk to a simulated customer, and call the right tools to search flights, change bookings, and issue refunds. The model doesn't need any domain knowledge; it's scored entirely on how well it executes tool calls across multi-step task trajectories. We run it continuously against the same provider endpoints that serve OpenRouter traffic, so a score reflects both the model and the provider running it. We use this benchmark because it has a high floor, so we can assess provider variance and not model capability. Our routing algorithm for tool call requests uses these same signals to send traffic to the best performing endpoints.

Last benchmark run Sep 12, 2026, 9:21 AM UTC

PaperGitHub

Which model leads the τ²-Bench Airline benchmark?

Google: Gemini 3.7 Flash leads τ²-Bench Airline at 80.6% as of Sep 12, 2026, 9:21 AM UTC. Airline tasks, policy, simulator and customer setup, step budget, and grading are held fixed across production provider endpoints, so the displayed results compare models and the providers serving them under the same evaluation conditions.

Tool-call errors before → after Auto Exacto

3.9% → 5.0%

Tool-call error rate on models enrolled in Auto Exacto, OpenRouter's automatic provider optimization for tool-calling requests.
Model comparisonCost efficiencyTool-call reliabilityLeaderboardWhy we run itWhat scores tell youHow tasks are scoredMethodologyAPI accessFAQ

Model comparison

Most Accurate

Favicon for google
Google: Gemini 3.7 Flash

80.6%

Best Value

Favicon for stepfun
StepFun: Step 3.7 Flash

$0.020/task

Fastest

Favicon for amazon
Amazon: Nova Micro 1.0

1.7m

Accuracy
Representative-run accuracy, best first.
Cost per task
Average cost per graded task, cheapest first.
Time per task
Average wall-clock time per task, fastest first; agents that loop or stall run long.

Cost efficiency

Accuracy vs. cost (Pareto frontier)
One point per model, using default routing (not pinned to a provider) when available. The line is the Pareto frontier: no model beats these on both accuracy and cost.

Tool-call reliability

Tool-call errors
Share of this benchmark's own requests where the model called a tool that doesn't exist, passed arguments that don't match the tool's schema, or emitted arguments that aren't valid JSON.

Leaderboard

Top-level rows use default routing where available; click a row to expand provider-pinned results.

τ²-Bench Airline leaderboard: model accuracy, cost, time, and output tokens per task by provider
#ModelStd dev
1
Google: Gemini 3.7 Flash
Pareto
80.6%--$0.0772.0m14.7k
2
Anthropic: Claude Fable 5
80.2%±3.0pp$0.962.2m5.79k
3
Claude Opus 5
79.6%±2.3pp$0.502.2m7.26k
4
Amazon: Nova Micro 1.0
78.7%±2.0pp$1.281.7m5.69k
5
Qwen: Qwen3.5 397B A17B
78.5%±3.6pp$0.0996.1m15.8k
6
Z.ai: GLM 5.3
78.3%±1.7pp$0.0983.1m9.09k
7
Anthropic: Claude Fable 5.1
78.0%--$0.692.0m6.91k
8
DeepSeek: DeepSeek V4 Pro 0813
77.7%±1.1pp$0.115.5m19.1k
9
Google: Gemini 3 Flash Preview
77.3%±2.0pp$0.122.5m20.5k
10
StepFun: Step 3.7 Flash
Pareto
77.3%--$0.0203.8m10.9k
11
Qwen: Qwen3.5-122B-A10B
77.1%±3.5pp$0.159.1m35k
12
Anthropic: Claude Opus 4.5
77.0%±2.6pp$0.532.9m10.4k
13
NVIDIA: Nemotron 3 Ultra
76.9%±0.8pp$0.102.6m8.57k
14
Qwen: Qwen3.8 27B
76.9%±3.0pp$0.0804.7m12.2k
15
Anthropic: Claude Opus 4.7
76.9%±1.5pp$0.411.7m5.68k
16
DeepSeek: DeepSeek V4.1 Flash
76.7%--$0.0283.7m15.4k
17
OpenAI: GPT-5.6 Sol Pro
76.7%--$1.865.6m22.9k
18
Qwen: Qwen3.8 2.4T A95B
76.7%±0.0pp$0.247.4m15.2k
19
Anthropic: Claude Sonnet 5
76.7%±2.0pp$0.202.3m7.91k
20
OpenAI: GPT-5.6 Sol
76.7%±0.7pp$0.212.1m5.14k
21
Anthropic: Claude Opus 4.8
76.2%±2.0pp$0.502.2m8.14k
22
Google: Gemma 4 31B
Pareto
76.1%±4.0pp$0.0165.5m8.28k
23
DeepSeek: DeepSeek V4 Pro 0423
76.0%±2.8pp$0.0422.7m7.59k
24
SpaceXAI: Grok 4.6
76.0%--$0.354.1m11.6k
25
Anthropic: Claude Opus 4.6
75.7%±2.5pp$0.483.3m10k
26
Google: Gemini 3.5 Flash Lite
75.7%±1.7pp$0.111.7m25.1k
27
OpenAI: GPT-5.1
75.7%±4.3pp$0.438.7m37.6k
28
OpenAI: GPT-5.4
75.7%±0.3pp$0.302.5m10.8k
29
Anthropic: Claude Sonnet 4.6
75.3%±1.7pp$0.323.3m11.4k
30
OpenAI: GPT-5.5
75.3%±2.7pp$0.502.6m8.28k
31
Google: Gemini 3.1 Pro Preview
75.3%±2.0pp$0.362.3m12.8k
32
Z.ai: GLM 5.3 Flash
Pareto
75.3%±2.0pp$0.0062.2m4.18k
33
Z.ai: GLM 5
75.3%±4.2pp$0.0362.0m5.54k
34
DeepSeek: DeepSeek V4 Flash 0423
75.1%±3.3pp$0.0092.4m8.29k
35
Z.ai: GLM 5.2
74.9%±2.7pp$0.0382.8m7.71k
36
Google: Gemini 3.5 Flash
74.7%±0.7pp$0.422.8m28.3k
37
Qwen: Qwen3.6 27B
74.1%±4.8pp$0.109.8m23.6k
38
Z.ai: GLM 5.1
74.1%±3.6pp$0.0572.4m5.73k
39
Xiaomi: MiMo-V2.5-Pro
74.1%±4.0pp$0.0286.6m12.7k
40
DeepSeek: DeepSeek V4 Flash Vision Exp
74.0%±0.0pp$0.0323.0m17.6k
41
MoonshotAI: Kimi K2.6
73.7%±2.7pp$0.0613.8m9.41k
42
OpenAI: GPT-5.6 Terra
73.3%±1.3pp$0.1001.6m4.17k
43
Meta: Muse Glimmer 30B
73.3%±3.5pp$0.0403.5m13.4k
44
OpenAI: GPT-5.2
73.3%--$0.243.5m12.5k
45
Qwen: Qwen3.5-35B-A3B
73.2%±8.2pp$0.0557.6m35k
46
Z.ai: GLM 4.7
73.0%±4.4pp$0.0372.6m5.62k
47
Anthropic: Claude Sonnet 4.5
73.0%±3.3pp$0.333.4m10.3k
48
Google: Gemini 3.6 Flash
73.0%±0.3pp$0.242.4m22.9k
49
DeepSeek: DeepSeek V4 Flash 0731
72.8%±5.9pp$0.0085.7m16.3k
50
Google: Gemini 3.1 Flash Lite
72.3%±1.0pp$0.0972.4m51.2k
51
DeepSeek: DeepSeek V3.2
72.1%±2.6pp$0.0234.6m8.07k
52
MoonshotAI: Kimi K2.7 Code
71.8%±3.3pp$0.0674.2m7.51k
53
OpenAI: GPT-5.6 Luna Pro
71.3%±1.3pp$0.0645.8m34.2k
54
MoonshotAI: Kimi K2.5
71.2%±2.7pp$0.0293.2m6k
55
MiniMax: MiniMax M2.7
71.1%±4.3pp$0.0182.3m5.87k
56
Qwen: Qwen3.6 35B A3B
71.1%±7.9pp$0.0433.7m24.1k
57
MoonshotAI: Kimi K3
70.9%±2.0pp$0.348.1m9.11k
58
MiniMax: MiniMax M3
70.9%±6.4pp$0.0232.2m6k
59
Z.ai: GLM 4.5 Air
70.9%±3.6pp$0.0142.4m4.53k
60
Auto Router (Beta)
70.7%--$0.242.5m17.9k
61
OpenAI: GPT-5.6 Luna
70.7%±0.0pp$0.0091.7m5.4k
62
Xiaomi: MiMo-V2.5
70.7%±2.4pp$0.0084.8m14.3k
63
OpenAI: GPT-5 Mini
69.7%±3.0pp$0.1311.6m58.4k
64
MoonshotAI: Kimi K2 Thinking
69.2%±10.8pp$0.0441.8m3.91k
65
DeepSeek: DeepSeek V3.1 Terminus
68.8%±4.1pp$0.0396.8m11.8k
66
Xiaomi: MiMo-V2-Flash
Pareto
68.5%±6.1pp$0.00643s3.58k
67
OpenAI: GPT-5
68.4%±3.8pp$0.6813.4m61.2k
68
Google: Gemma 4 26B A4B
68.3%±3.9pp$0.0174.8m13k
69
Anthropic: Claude Sonnet 4
68.0%±2.0pp$0.332.9m8.95k
70
DeepSeek: DeepSeek V3.2 Exp
67.8%±4.3pp$0.0567.7m8.93k
71
Anthropic: Claude Haiku 4.5
67.8%±3.3pp$0.122.0m12.7k
72
Qwen: Qwen3.5-9B
67.5%±5.0pp$0.0249.6m31.3k
73
Thinking Machines: Inkling
67.3%±3.6pp$0.08062s3.68k
74
OpenAI: GPT-5.2 Chat
67.3%--$0.1180s4.28k
75
OpenAI: GPT-5.4 Nano
67.0%±0.3pp$0.0332.5m17.5k
76
Z.ai: GLM 4.6
66.6%±11.2pp$0.0363.4m6.05k
77
MiniMax: MiniMax M2.1
66.6%±3.6pp$0.01582s5.49k
78
OpenAI: gpt-oss-120b
64.1%±3.4pp$0.0133.5m13.3k
79
MiniMax: MiniMax M2.5
64.1%±4.8pp$0.0162.8m6.61k
80
OpenAI: GPT-5.4 Mini
63.0%±5.0pp$0.0922.2m12.5k
81
NVIDIA: Nemotron 3.5 Lightning
62.9%--$0.0202.7m22.2k
82
Z.ai: GLM 4.7 Flash
61.9%±6.3pp$0.0082.6m7.86k
83
Google: Gemini 2.5 Pro
61.3%±0.7pp$0.223.2m15.1k
84
Qwen: Qwen3 235B A22B Thinking 2507
61.3%±4.1pp$0.0586.8m20.6k
85
Thinking Machines: Inkling Small
60.9%±3.1pp$0.0473.8m3.8k
86
DeepSeek: DeepSeek V3.1
60.2%±10.5pp$0.0595.9m9.02k
87
Qwen: Qwen3 Coder Next
58.4%±6.5pp$0.02461s2.79k
88
Google: Gemini 2.5 Flash
57.7%±0.3pp$0.02979s8.17k
89
inclusionAI: Ling 3.0 Flash
57.3%--$0.0122.0m16.8k
90
MoonshotAI: Kimi K2 0905
54.9%±5.2pp$0.112.6m3.13k
91
DeepSeek: R1 0528
54.5%±5.8pp$0.08413.4m17.3k
92
NVIDIA: Nemotron 3 Nano 30B A3B
52.3%±2.6pp$0.0247.5m50k
93
OpenAI: gpt-oss-20b
51.4%±2.5pp$0.02318.7m96.5k
94
OpenAI: GPT-4.1
48.7%±0.7pp$0.1251s2.66k
95
OpenAI: GPT-5 Nano
47.6%±4.5pp$0.04311.6m99.7k
96
Google: Gemini 2.5 Flash Lite
47.3%--$0.0242.9m45.1k
97
Qwen: Qwen3 Next 80B A3B Instruct
47.0%±3.6pp$0.01562s2.3k
98
Qwen: Qwen3 Coder 480B A35B
46.7%±5.2pp$0.04865s1.89k
99
OpenAI: GPT-4o
46.3%±1.7pp$0.2052s2.08k
100
OpenAI: GPT-5.3 Chat
46.0%--$0.09965s2.75k
101
Qwen: Qwen3 235B A22B Instruct 2507
45.9%±4.1pp$0.0181.7m2.28k
102
OpenAI: GPT-4o (2024-08-06)
44.7%±1.3pp$0.2147s2.22k
103
Meta: Llama 4 Maverick
44.6%±2.3pp$0.0411.8m1.94k
104
OpenAI: GPT-4.1 Mini
44.0%±1.3pp$0.02974s2.45k
105
Qwen: Qwen3 30B A3B
43.7%±6.1pp$0.0153.1m10.3k
106
Mistral: Mistral Small 4
43.7%±2.3pp$0.0141.6m10.2k
107
Qwen: Qwen3 Coder 30B A3B Instruct
43.1%±0.6pp$0.0243.6m3.15k
108
DeepSeek: DeepSeek V3 0324
43.0%±3.8pp$0.0425.7m8.59k
109
Qwen: Qwen3 14B
42.3%±1.0pp$0.0308.7m20.1k
110
Qwen: Qwen3 32B
42.1%±9.9pp$0.0157.7m13.2k
111
Qwen: Qwen3 VL 235B A22B Instruct
40.7%±4.4pp$0.0352.3m3.44k
112
OpenAI: GPT-4o (2024-05-13)
40.0%--$1.1372s3.25k
113
Meta: Llama 3.3 70B Instruct
38.7%±3.6pp$0.01243s550
114
DeepSeek: DeepSeek V3
38.3%±0.3pp$0.0714.7m7.51k
115
Qwen: Qwen3 30B A3B Instruct 2507
37.7%±3.4pp$0.0161.9m2.64k
116
Qwen: Qwen3 VL 30B A3B Instruct
34.7%±3.2pp$0.0733.6m6.68k
117
Qwen: Qwen3 VL 8B Instruct
32.8%±4.5pp$0.02067s2.67k
118
Meta: Llama 3.1 8B Instruct
31.4%±6.5pp$0.00665s1.52k
119
OpenAI: GPT-4o-mini
27.0%±3.0pp$0.02166s3.95k
120
Qwen2.5 72B Instruct
26.0%±1.3pp$0.0799.8m5.67k
121
Mistral: Mistral Nemo
18.7%±6.3pp$0.0201.8m2.7k
122
Qwen: Qwen2.5 7B Instruct
16.7%--$0.01758s2.49k
123
OpenAI: GPT-4.1 Nano
10.3%±0.3pp$0.00761s2.32k

Why we run this benchmark

It's a tool-calling benchmark that is hard to game. Grading depends on live tool-call trajectories rather than memorized answers, so it resists training-data leakage better than Q&A-style evals. It exercises every tool-calling failure mode (wrong arguments, skipped policy checks, giving up, hallucinated confirmations) at a relatively low cost per run. The relative scores also carry more signal than the absolute ones. The same model can score differently across providers, and those deltas are what Exacto routing uses to pick higher-accuracy endpoints.

Each task is a simulated airline support conversation with a scripted user, a toolbox (flight search, booking changes, refunds, loyalty policies), and a gold reference solution. A task passes only if the final database state and the messages to the user match the reference; partial credit is not awarded.

What the scores can and can't tell you

There is still headroom. Top models fail roughly one in five tasks, and the airline domain is the hardest τ²-Bench split. Accuracy differences here separate models that follow multi-step policies from ones that merely chat well.

The floor is high, though. Many tasks reward inaction. A refusal task with an empty gold action list passes for any agent that changes nothing. Even weak models score well above zero, so the meaningful spread sits at the top of the range.

The benchmark is public, so tasks may appear in training corpora. Contamination inflates scores less here than in Q&A-style evals, though, since a leaked task still has to be executed correctly, step by step, against a live database.

Scoring fidelity has limits. The checker verifies two things: the final database hash and exact substring matches in the agent's messages. Each task's natural-language assertions ("agent should refuse the cancellation") are metadata, and no judge model reads the transcript. So a savings calculation fails if the agent says "$23,552.50" when the checker greps for "23553".

The user simulator matters too. We pin it to gemini-2.5-flash so agent scores stay comparable, but the sim is itself an LLM with failure modes of its own. It can stop the conversation before the agent finishes, leak its hidden task instructions, or keep a stuck agent looping until the 200-step ceiling kills the run. Swapping the sim model shifts absolute scores, which is why cross-paper τ²-Bench numbers rarely line up exactly.

How a task is scored

Every task ships a gold solution: a list of tool calls, strings the agent must say, and natural-language assertions. After the conversation ends, the checker replays the gold tool calls against a fresh database and compares hashes with the agent's final database. It then greps the agent's messages for each required string. The reward is the product of those two checks:

reward = db_match × communicate_met   // each ∈ {0, 1}
db_match        = hash(agent DB) == hash(gold DB)
communicate_met = every required string appears in an agent message
any run that hits MAX_STEPS instead of a clean stop scores 0 outright

The rollouts below are from real runs, with gemini-2.5-flash as the user simulator throughout.

reward = 1

Pass: three changes in one request, all three land

Task 17 agent: openai/gpt-5.1

For reservation FQ8APE: add 3 checked bags, swap the passenger to Omar Rossi, and upgrade basic economy to economy, paying with a gift card.

  • Database must match the gold state: update_reservation_flights (economy upgrade), update_reservation_passengers, and update_reservation_baggages with exact arguments
  • communicate_info is empty, so no string check applies
db ✓communicate ✓USER_STOP

The agent looked up the user, found the right reservation among several, confirmed the changes and payment method, then made all three writes: passenger swap, cabin upgrade, and bags. The final database hashes match the gold state and the run ends on USER_STOP, so reward is 1. This is what the eval is designed to measure: multi-step tool use under policy constraints, done correctly.

Methodology

Scores aggregate all successful runs, weighted by task count, with a minimum of 45 graded tasks per model-provider pair. A model's headline score uses its default routing (not pinned to a provider) when one exists; otherwise it falls back to the median provider. The standard deviation is measured across runs for that representative result. Cost, time, and token figures are per-task averages from the same runs. Best value is the cheapest Pareto-optimal model within 5 points of the top score.

These are the same measurements that power Exacto routing. See the docs for how routing works, or browse all models to try one.

API access

These scores are available through OpenRouter's public benchmarks API, so you can retrieve the same model-level results programmatically.

GET https://openrouter.ai/api/v1/benchmarks?source=openrouter
Authorization: Bearer <API key>

Use task_type=agentic to filter to tau_bench_verified_airline. Each item represents one model and includes accuracy, accuracy_stddev, avg_cost_per_task, total_tasks, and last_run_timestamp. See the benchmarks API docs.

Frequently asked questions

It tests whether a model can do an airline support agent’s job: follow a policy manual, talk to a simulated customer, and call the right tools to search flights, change bookings, and issue refunds. The model needs no domain knowledge, and is scored entirely on how well it executes tool calls across multi-step task trajectories.
A task is graded pass or fail. The agent has to finish the conversation within its step budget, leave the booking system in the state the reference solution produces, and tell the customer what it did; missing any one of those scores zero, and partial credit is not awarded.
Runs execute against the same provider endpoints that serve OpenRouter traffic, so a score reflects both the model and the provider running it. These accuracy scores are one of the signals Exacto routing uses to steer tool-calling traffic away from endpoints that fall behind their peers.
The customer side of every conversation is played by another model, and OpenRouter keeps that simulator fixed so agent scores stay comparable across runs. The simulator has failure modes of its own, so absolute scores shift whenever it changes, which is why τ²-Bench numbers from different sources rarely line up exactly.
Scores aggregate repeated runs of the same task set, weighted by how many tasks each run graded. A model’s headline score is a single representative result rather than its best-performing provider.
Yes. The public benchmarks API returns the same model-level results from GET https://openrouter.ai/api/v1/benchmarks?source=openrouter with an API key. Filter with task_type=agentic to reach tau_bench_verified_airline.
OpenRouterOpenRouter
© 2026 OpenRouter, Inc

Product

  • Chat
  • Rankings
  • Benchmarks
  • Apps
  • Discover
  • Models
  • Collections
  • Providers
  • Pricing
  • Business
  • Enterprise
  • Labs

Company

  • About
  • Blog
  • Careers
    Hiring
  • Privacy
  • Terms of Service
  • Trust Center
  • Support
  • Works With OR
  • Data
  • Brand

Developer

  • Documentation
  • API Reference
  • Developer Platform
  • Status

Connect

  • Discord
  • GitHub
  • LinkedIn
  • X
  • YouTube