Atlas

Benchmarks

← All benchmarks

CL-bench (Overall)

General QA · 2026-02-03

CL-bench evaluates long-context learning across domain-knowledge reasoning, rule-system application, procedural execution, and empirical discovery tasks. Scores are reported as overall percentage-style performance, with higher values indicating stronger context learning.

Top models (higher is better)

ModelScore
GPT-5.427.9
GPT-5.123.7
Grok 4.2022.2
Opus 4.521.1
Gemini 3.1 Pro Preview20.8
Opus 4.620.7
Qwen3.6 Plus (2026-04-02)20.3
Qwen3.5 Plus (2026-02-15)19.8
Kimi K2.519.3
GLM-518.7
GPT-5.218.2
o317.8
GLM-4.715.9
Gemini 3 Pro Preview15.8
MiMo-V2-Pro15.7
Qwen3-Max (2025-09-23)14.5
DeepSeek-V3.2-Exp13.2
DeepSeek-V3.212.4
MiniMax M2.511.4
Loading Atlas data…