Atlas

Benchmarks

← All benchmarks

ComplexConstraints

Chat & Writing · 2026-06-03

Instruction-following benchmark for interlocking, conditional constraints in professional work.

Top models (higher is better)

ModelScore
GPT-5.6 Sol50.5
GPT-5.549.5
GPT-5.446.0
Gemini 3.1 Pro Preview43.7
Gemini 3.6 Flash40.0
Claude Fable 538.1
Muse Spark 1.138.1
Kimi K337.9
Claude Opus 537.3
Gemini 3.5 Flash37.1
Opus 4.636.3
Grok 4.535.9
Opus 4.835.6
Sonnet 4.634.0
Qwen3.7-Max33.5
Opus 4.731.6
Kimi K2.630.1
GLM-5.229.3
DeepSeek-V4-Pro28.0
Gemini 3.5 Flash-Lite23.2
DeepSeek-V4-Flash20.9
Kimi K2.518.2
Nemotron 3 Ultra 550B A55B17.8
Qwen3.5 Plus (2026-02-15)17.8
Grok 4.2016.4
ERNIE 5.114.7
DeepSeek-V3.22.2
Mistral Large 3 675B Instruct 25120.4
Inkling0.3
ERNIE 4.5 300B A47B0.0
Nova 2 Pro (Preview)0.0
Loading Atlas data…