CL-bench Life (Overall)
General QA · 2026-04-29
CL-bench Life evaluates long-context learning in life-scenario tasks involving social communication, fragmented revisions, and behavioral records. Scores are reported as overall percentage-style performance, with higher values indicating stronger context learning.
Top models (higher is better)
| Model | Score |
|---|---|
| GPT-5.5 | 22.2 |
| GPT-5.4 | 21.7 |
| GPT-5.1 | 17.3 |
| Opus 4.6 | 17.0 |
| Gemini 3.1 Pro Preview | 16.9 |
| DeepSeek-V4-Pro | 13.5 |
| Kimi K2.5 | 13.2 |
| DeepSeek-V3.2-Exp | 9.5 |
| DeepSeek-V3.2 | 7.4 |
| MiniMax M2.5 | 6.3 |