Atlas

Benchmarks

← All benchmarks

SOLO Bench Easy

Chat & Writing · 2025-05-01

SOLO Bench Easy asks for 250 unique constrained sentences and reports the percentage score from the rule-based evaluator. Higher scores indicate better compliance with the vocabulary, uniqueness, and sentence-format constraints.

Top models (higher is better)

ModelScore
Gemini 2.5 Pro Preview 03-2574.8
o356.4
Claude 3.7 Sonnet34.0
Grok 331.2
DeepSeek-R128.4
GPT-4.526.8
DeepSeek-V3-032420.0
Gemini 2.5 Flash Preview 04-1716.8
GPT-4.19.2
Qwen3-235B-A22B8.4
Llama 3.1 Nemotron Ultra 253B V18.0
Qwen3 32B5.2
Qwen2.5-VL-72B-Instruct5.2
Llama 3.1 405B Instruct4.4
Llama 4 Maverick Instruct4.0
Gemma 3 27B IT1.2
Llama 3.3 70B Instruct0.4
Gemma 3 4B IT0.0
Qwen3 8B0.0
O4 Mini0.0
Llama 4 Scout Instruct0.0
Loading Atlas data…