AA-LCR v1.1
Search
Artificial Analysis Long Context Reasoning v1.1: 100 questions over long documents, evaluated with three attempts per question. The public model-data schema identifies the lcr_1_1 series; earlier unversioned observations remain separate.
Top models (higher is better)
| Model | Score |
|---|---|
| Kimi K3 | 88.7 |
| Step 5 Preview | 88.3 |
| MiMo-V2.6-Pro | 86.3 |
| Claude Fable 5.1 | 85.3 |
| Claude Opus 5.5 | 84.7 |
| GPT-5.5 | 84.3 |
| DeepSeek V4.1 Flash | 84.0 |
| Gemini 3.8 Flash | 84.0 |
| GPT-5.6 Sol | 84.0 |
| GPT-5.6 Luna | 83.7 |
| GPT-6 Sol | 83.7 |
| GPT-5.3-Codex | 83.3 |
| GPT-6 Luna | 83.3 |
| Muse Glimmer 30B | 83.3 |
| Agnes 2.5 Pro Beta | 83.0 |
| Gemini 3.7 Flash | 83.0 |
| GPT-5.6 Terra | 83.0 |
| MiniMax M3 | 83.0 |
| Muse Spark 1.3 | 83.0 |
| GPT-5.2 | 82.7 |
| Claude Fable 5 | 82.3 |
| GPT-5.2-Codex | 82.3 |
| Claude Opus 5 | 82.0 |
| Sonnet 5 | 82.0 |
| Gemini 3.1 Pro Preview | 82.0 |
| GPT-5.4 | 82.0 |
| Qwen3.8 27B | 82.0 |
| Nex-N2-Pro | 81.7 |
| DeepSeek V4 Flash Vision Exp | 81.3 |
| Agnes 3.0 Flash (hosted) | 81.0 |
| Grok 4.6 | 81.0 |
| Kimi K2.6 | 81.0 |
| GPT-6 Astra | 80.7 |
| Qwen3.6-Max-Preview | 80.7 |
| DeepSeek V4 Pro 0813 | 80.3 |
| Qwen3.8 2.4T A95B | 80.3 |
| Qwen3.8 Max (0902) | 80.3 |
| Sonnet 4.6 | 80.0 |
| Gemini 3.6 Flash | 80.0 |
| GLM-5.3 Flash | 80.0 |