Atlas

Benchmarks

← All benchmarks

LMArena Agent Arena Steering Burden

Agents

Signed percentage-point Agent Arena steering_burden diagnostic, displayed by the source as Steerability. Measures response to user corrections. Kept separate from the earlier steerability source signal because its key and observation population changed on the September 30 snapshot.

Top models (higher is better)

ModelScore
Gemini 4 Argon15.9
GPT-6 Sol11.9
Claude Fable 59.0
Claude Fable 5.19.0
GPT-6 Astra8.2
Opus 4.87.5
Claude Opus 57.0
GPT-5.6 Sol6.2
Muse Spark 1.35.5
GPT-5.45.4
Sonnet 55.4
GPT-5.53.9
Gemini 3.8 Flash3.3
Qwen3.8-Max3.1
GPT-5.6 Terra2.5
Grok 4.52.1
GLM-5.21.8
Kimi K31.4
DeepSeek V4.1 Flash1.2
Grok 4.61.0
Hy4 Preview0.9
GLM-5.30.9
Grok 4.70.4
Hy30.4
DeepSeek V4 Pro 08130.3
GPT-5.6 Luna0.1
GLM-5.3 Flash-0.2
Qwen3.8 27B-0.4
Qwen3.8-Flash-Next-0.8
Gemini 3.7 Flash-1.0
Gemini 3.1 Pro Preview-4.5
Muse Spark 1.2-4.9
Muse Spark 1.1-6.2
Qwen3.7-Max-6.4
MiniMax M3-6.9
MiMo-V2.5-Pro-7.4
Gemini 3.6 Flash-7.5
Qwen3.7-Plus-8.4
Solar Pro 4-9.1
Inkling-11.9
Loading Atlas data…