Atlas

Benchmarks

← All benchmarks

Judgemark v4

Chat & Writing · 2026-05-04

Judgemark v4 tests how well an LLM judge separates stronger from weaker fixed creative-writing samples. Its 0–100 score combines writing-score separability metrics.

Top models (higher is better)

ModelScore
Opus 4.690.7
GPT-5.587.8
Opus 4.784.0
Sonnet 4.682.1
Claude Opus 578.8
Gemini 3.1 Pro Preview78.7
Opus 4.878.0
Grok 4.577.1
GLM-5.273.2
Gemma 4 31B IT72.3
GPT-5.472.1
Sonnet 570.9
Gemini 3.5 Flash68.8
GLM-5.167.2
Qwen3.5-27B60.5
Gemini 3.1 Flash-Lite Preview58.8
MiMo-V2.5-Pro57.9
Kimi K2.657.3
Qwen3.5 397B A17B54.2
Qwen3.6 27B53.4
Gemma 4 26B A4B IT53.0
Grok 4.349.6
DeepSeek-V4-Pro47.1
Gemini 3 Flash Preview46.1
Qwen3.6-Max-Preview45.1
Qwen3.5 35B A3B40.5
DeepSeek-V4-Flash36.8
Qwen3.6 35B A3B32.7
Qwen3.5-9B32.4
Qwen3.6-Flash-2026-04-1631.2
Step 3.7 Flash28.0
GPT-5.4 Nano27.5
GPT-5.4 Mini27.3
Mistral Small 416.5
gpt-oss-120b15.5
Loading Atlas data…