Atlas

Benchmarks

← All benchmarks

τ²-bench (Retail)

Agents · 2025-06-09

The retail domain of τ²-bench evaluates an agent handling customer-service tasks against a policy document and a set of tools, with a simulated user. Scores report end-to-end task success as pass^1 over independent trials.

Top models (higher is better)

ModelScore
Opus 4.691.9
Sonnet 4.691.7
Gemini 3.1 Pro Preview90.8
Qwen3.5 397B A17B84.4
GPT-5.282.0
Opus 4.579.6
Gemini 3 Flash Preview76.8
GLM-573.7
Sonnet 4.572.4
Loading Atlas data…