Atlas

Benchmarks

← All benchmarks

CL-bench Life (Overall)

General QA · 2026-04-29

CL-bench Life evaluates long-context learning in life-scenario tasks involving social communication, fragmented revisions, and behavioral records. Scores are reported as overall percentage-style performance, with higher values indicating stronger context learning.

Top models (higher is better)

ModelScore
GPT-5.522.2
GPT-5.421.7
GPT-5.117.3
Opus 4.617.0
Gemini 3.1 Pro Preview16.9
DeepSeek-V4-Pro13.5
Kimi K2.513.2
DeepSeek-V3.2-Exp9.5
DeepSeek-V3.27.4
MiniMax M2.56.3
Loading Atlas data…