Atlas

Benchmarks

← All benchmarks

Creative Writing v2

Chat & Writing · 2024-04-20

Creative Writing v2 scores model responses to 24 creative-writing prompts, repeated over 10 iterations, against 27 narrowly defined rubric criteria applied by an LLM judge. Scoring is absolute rubric grading rather than comparative: EQ-Bench moved to pairwise Elo only in v3.

Top models (higher is better)

ModelScore
DeepSeek-R187.1
Gemini 2.5 Pro Experimental 03-2586.2
Gemma 3 27B IT85.3
QwQ-32B84.7
GPT-4o (2024-11-20)84.6
DeepSeek-V3-032484.5
Claude 3.7 Sonnet84.5
Claude 3.5 Sonnet (Oct 2024)83.0
ChatGPT-4o Latest (September 2024 benchmark entry)82.5
Gemini 2.0 Flash Thinking Experimental 121982.5
GPT-4.581.7
Gemini 1.5 Pro 00281.4
Gemini 2.0 Flash 00181.4
DeepSeek-V381.2
o1 Preview80.5
Gemini 1.5 Pro Experimental 080180.3
Gemini 1.5 Pro 00180.3
LFM-7B79.7
WizardLM-2 8x22B78.9
Claude 3.5 Sonnet (June 2024)78.8
GPT-4o Mini78.4
Grok Beta78.1
Mistral NeMo Instruct 240777.5
GPT-4 0125 Preview77.4
Gemma 2 27B IT77.2
Mistral Large 2 (Instruct 2407)77.2
o1-mini76.3
DeepSeek-R1-Distill-Llama-70B76.2
Gemma 2 9B IT76.2
C4AI Command R+ 08-202476.1
GPT-4o75.6
Qwen1.5 110B Chat75.3
Opus 373.6
DeepSeek-R1-Distill-Qwen-32B73.0
Mistral Small Instruct 240972.4
Qwen2.5 72B Instruct72.2
Llama 3.1 405B Instruct72.0
Gemini 1.5 Flash72.0
Llama 3 70B Instruct71.3
Yi 34B Chat71.1
Loading Atlas data…