CL-bench (Overall)
General QA · 2026-02-03
CL-bench evaluates long-context learning across domain-knowledge reasoning, rule-system application, procedural execution, and empirical discovery tasks. Scores are reported as overall percentage-style performance, with higher values indicating stronger context learning.
Top models (higher is better)
| Model | Score |
|---|---|
| GPT-5.4 | 27.9 |
| GPT-5.1 | 23.7 |
| Grok 4.20 | 22.2 |
| Opus 4.5 | 21.1 |
| Gemini 3.1 Pro Preview | 20.8 |
| Opus 4.6 | 20.7 |
| Qwen3.6 Plus (2026-04-02) | 20.3 |
| Qwen3.5 Plus (2026-02-15) | 19.8 |
| Kimi K2.5 | 19.3 |
| GLM-5 | 18.7 |
| GPT-5.2 | 18.2 |
| o3 | 17.8 |
| GLM-4.7 | 15.9 |
| Gemini 3 Pro Preview | 15.8 |
| MiMo-V2-Pro | 15.7 |
| Qwen3-Max (2025-09-23) | 14.5 |
| DeepSeek-V3.2-Exp | 13.2 |
| DeepSeek-V3.2 | 12.4 |
| MiniMax M2.5 | 11.4 |