Atlas

Benchmarks

← All benchmarks

MLCR-AA Conciseness

Professional Work

Share of MLCR-AA responses meeting the length and answer-presence requirements. This does not measure answer correctness and is excluded from scalar capability fitting.

Top models (higher is better)

ModelScore
GPT-6 Sol100.0
Nova Lite100.0
Command A+ (May 2026, BF16)99.4
Gemini 3.1 Pro Preview99.4
GPT-5.4 Mini99.4
Grok 4.699.4
Nemotron 3 Super 120B A12B99.4
Gemini 3.1 Flash-Lite Preview98.9
Gemini 3 Flash Preview98.9
Grok 4.398.9
Qwen3.7-Plus98.9
Gemini 2.5 Flash-Lite98.3
Gemini 3.5 Flash-Lite98.3
Gemini 3.6 Flash98.3
Gemini 3.7 Flash98.3
Grok 4.598.3
Llama 4 Maverick Instruct98.3
Opus 4.797.8
Gemini 2.5 Flash97.8
Gemini 3.5 Flash97.8
GPT-4o Mini97.8
GPT-5.6 Sol97.8
GPT-6 Luna97.8
Grok 4.797.8
Kimi K2.7 Code97.8
Nemotron 3 Ultra 550B A55B97.8
DeepSeek-V4-Flash-073197.2
DeepSeek-V4-Flash97.2
Gemini 3.8 Flash97.2
Nemotron 3.5 Lightning 30B A3B (precision unspecified)97.2
DeepSeek V4.1 Flash96.7
GLM-5.3 Flash96.7
GPT-5.4 Nano96.7
GPT-5.596.7
GPT-5.6 Luna96.7
GPT-5.6 Terra96.7
GPT-6 Astra96.7
Inkling96.7
Step 5 Preview96.7
Opus 4.896.1
Loading Atlas data…