Atlas

Benchmarks

← All benchmarks

Instruction Following

Chat & Writing · 2024-05-29

Score on Scale SEAL's precise instruction-following prompts.

Top models (higher is better)

ModelScore
O192.0
DeepSeek-R187.8
Gemini 2.0 Flash Experimental86.6
o1 Preview86.6
Claude 3.5 Sonnet (June 2024)86.0
GPT-4o85.3
Llama 3.1 405B Instruct84.8
Gemini 1.5 Pro Experimental 082784.2
GPT-4 0125 Preview83.2
Mistral Large 2 (Instruct 2407)82.8
GPT-4o (2024-11-20)82.5
DeepSeek-V382.3
Llama 3.2 90B Vision Instruct82.1
Llama 3 70B Instruct81.2
GPT-4o (2024-08-06)80.2
Opus 380.1
Mistral Large 1.079.9
GPT-4 Turbo (1106 Preview)79.5
Gemini 1.5 Pro Preview 051479.4
Loading Atlas data…