Atlas

Benchmarks

← All benchmarks

AA-LCR v1.1

Search

Artificial Analysis Long Context Reasoning v1.1: 100 questions over long documents, evaluated with three attempts per question. The public model-data schema identifies the lcr_1_1 series; earlier unversioned observations remain separate.

Top models (higher is better)

ModelScore
Kimi K388.7
Step 5 Preview88.3
MiMo-V2.6-Pro86.3
Claude Fable 5.185.3
Claude Opus 5.584.7
GPT-5.584.3
DeepSeek V4.1 Flash84.0
Gemini 3.8 Flash84.0
GPT-5.6 Sol84.0
GPT-5.6 Luna83.7
GPT-6 Sol83.7
GPT-5.3-Codex83.3
GPT-6 Luna83.3
Muse Glimmer 30B83.3
Agnes 2.5 Pro Beta83.0
Gemini 3.7 Flash83.0
GPT-5.6 Terra83.0
MiniMax M383.0
Muse Spark 1.383.0
GPT-5.282.7
Claude Fable 582.3
GPT-5.2-Codex82.3
Claude Opus 582.0
Sonnet 582.0
Gemini 3.1 Pro Preview82.0
GPT-5.482.0
Qwen3.8 27B82.0
Nex-N2-Pro81.7
DeepSeek V4 Flash Vision Exp81.3
Agnes 3.0 Flash (hosted)81.0
Grok 4.681.0
Kimi K2.681.0
GPT-6 Astra80.7
Qwen3.6-Max-Preview80.7
DeepSeek V4 Pro 081380.3
Qwen3.8 2.4T A95B80.3
Qwen3.8 Max (0902)80.3
Sonnet 4.680.0
Gemini 3.6 Flash80.0
GLM-5.3 Flash80.0
Loading Atlas data…