Atlas

Benchmarks

← All benchmarks

LLM Creative Story Writing

Chat & Writing · 2025-01-05

Independent creative-writing benchmark comparing short stories generated by LLMs. The current leaderboard reports model scores and win probabilities from pairwise comparisons.

Top models (higher is better)

ModelScore
Claude Fable 53.3
GPT-5.53.1
GPT-5.6 Sol3.0
GPT-5.42.8
Opus 4.72.5
Sonnet 4.62.3
Opus 4.61.8
Opus 4.81.5
Muse Spark 1.11.4
GPT-5.21.0
GLM-5.21.0
Kimi K2.60.8
MiniMax M30.7
Mistral Medium 3.10.3
DeepSeek-V4-Pro0.2
MiMo-V2.5-Pro0.0
Qwen3 Max (rolling alias)0.0
Qwen3.6-Max-Preview-0.3
GLM-5.1-0.4
Kimi K2.5-0.5
ERNIE 5.1-0.6
MiMo-V2-Pro-0.6
Mistral Large 3 675B Instruct 2512-1.3
Gemma 4 31B IT-1.3
Gemini 3.5 Flash-1.4
Doubao Seed 2.0 Pro-1.5
Gemini 3.1 Pro Preview-1.8
Qwen3.6 Plus (2026-04-02)-1.8
Mistral Medium 3.5-2.0
Qwen3.7-Max-2.0
DeepSeek-V3.2-2.3
gpt-oss-120b-2.6
MiniMax M2.7-3.2
Grok 4.3-3.7
Grok 4.5-4.6
Loading Atlas data…