Atlas

Benchmarks

← All benchmarks

τ²-bench (Telecom)

Agents · 2025-06-09

The telecom domain of τ²-bench evaluates collaborative troubleshooting in its dual-control setup. Scores report end-to-end task success.

Top models (higher is better)

ModelScore
Gemini 3.1 Pro Preview99.3
GLM-5.299.1
JT-35B-Flash99.1
GPT-5.498.9
GLM 4.7 Flash98.8
GPT-5.298.7
Claude Fable 598.5
GLM 5 Turbo98.5
GLM 5V Turbo98.5
Step 3.7 Flash98.5
GLM-598.2
Gemini 3 Pro Preview98.0
GPT-5.598.0
Sonnet 4.697.9
GLM-5.197.7
Grok 4.397.7
Qwen3.6 Plus (2026-04-02)97.7
Grok 4.2096.5
DeepSeek-V4-Pro96.2
GLM-4.795.9
Kimi K2.595.9
Kimi K2.695.9
Qwen3.6-Max-Preview95.9
DeepSeek-V4-Flash95.6
Gemini 3.5 Flash95.6
Qwen3.5 397B A17B95.6
MiniMax M2.595.3
Qwen3.6 35B A3B95.3
MiMo-V2-Flash95.0
MiMo-V2-Pro95.0
Qwen3.7-Max94.7
Opus 4.894.4
Step 3.5 Flash94.4
MiMo-V2.5-Pro94.2
Mistral Medium 3.594.2
Qwen3.6 27B94.2
Qwen3.5-27B93.9
Qwen3.5 122B A10B93.6
GPT-5.4 Mini93.4
Grok 4.1 Fast93.3
Loading Atlas data…