Atlas

Benchmarks

← All benchmarks

LLM Writing Style Diversity

Chat & Writing · 2025-12-18

LLM Writing Style Diversity measures the within-model stylistic range of at least 400 flash-fiction stories using mean pairwise weighted Gower distance over fixed numeric and categorical style features. The score measures diversity rather than writing quality, with higher values indicating a broader observed range of styles.

Top models (higher is better)

ModelScore
GPT-50.2
GPT-5.20.2
GPT-5.10.2
GPT-5 Pro0.2
Kimi K2 Thinking0.2
GLM-4.50.2
Grok 4.1 Fast0.2
DeepSeek-V3.2-Exp0.2
Opus 4.10.2
Kimi K2 Instruct 09050.2
Kimi K2 Instruct0.2
GLM 4.60.2
Qwen3 235B A22B Thinking 25070.2
Sonnet 4.50.2
Mistral Large 3 675B Instruct 25120.2
o3-pro0.2
Grok 40.2
ERNIE 4.5 300B A47B0.2
Gemini 2.5 Pro0.2
Gemini 3 Pro Preview0.2
Mistral Medium 3.10.2
DeepSeek-V3.10.2
Opus 4.50.2
Llama 4 Maverick0.2
Qwen3 Max (rolling alias)0.2
Command A0.2
gpt-oss-120b0.2
Loading Atlas data…