Fiction.liveBench
Chat & Writing · 2025-02-19
Long-context creative-writing benchmark testing whether models understand and reason over extended fiction passages.
Top models (higher is better)
| Model | Score |
|---|---|
| GPT-5 | 97.2 |
| o3-pro | 97.2 |
| Grok 4 | 94.4 |
| Grok 4 Fast | 94.4 |
| Gemini 2.5 Pro Preview 06-05 | 91.7 |
| o3 | 88.9 |
| Kimi K2.5 | 86.1 |
| Claude 3.7 Sonnet | 83.3 |
| O1 | 83.3 |
| QwQ-32B | 83.3 |
| Gemini 2.5 Flash Preview 05-20 | 77.8 |
| O4 Mini | 77.8 |
| Qwen3 32B | 74.2 |
| DeepSeek-R1 | 69.4 |
| GPT-5 Mini | 69.4 |
| MiniMax M1 80k | 69.4 |
| Qwen3-235B-A22B | 67.7 |
| ChatGPT-4o Latest observed 2025-01-29 | 66.7 |
| Gemini 2.5 Pro Experimental 03-25 | 66.7 |
| Gemini 2.5 Pro Preview 05-06 | 66.7 |
| Gemini 2.5 Pro Preview 03-25 | 66.7 |
| Grok 3 Mini | 66.7 |
| Kimi K2 Instruct 0905 | 66.7 |
| Qwen2.5-Max | 66.7 |
| Qwen3-Max (2025-09-23) | 66.7 |
| GPT-4.1 | 63.9 |
| GPT-4.5 | 63.9 |
| Qwen3 14B | 62.5 |
| Qwen3 8B | 62.1 |
| Opus 4 | 61.1 |
| Gemini 2.0 Flash 001 | 61.1 |
| Kimi K2 Instruct | 61.1 |
| Grok 3 | 58.3 |
| DeepSeek-V3.1 | 52.8 |
| DeepSeek-V3.2-Exp | 52.8 |
| Gemini 2.0 Flash Thinking Experimental 01-21 | 52.8 |
| DeepSeek-V3-0324 | 50.0 |
| o3-mini | 50.0 |
| Gemini 2.5 Flash-Lite Preview 06-17 | 47.2 |
| Gemini 2.5 Flash Preview 04-17 | 47.2 |