Atlas

Benchmarks

← All benchmarks

Coding

Code · 2024-05-29

Elo rating on Scale SEAL's coding prompt set across programming tasks and languages.

Top models (higher is better)

ModelScore
o1-mini1237
o3-mini1137
GPT-4o (2024-11-20)1132
o1 Preview1123
Gemini 2.0 Flash Experimental1111
Gemini Exp-12061109
Gemini 2.0 Flash Thinking Experimental 01-211108
DeepSeek-R11100
O11083
Claude 3.5 Sonnet (June 2024)1079
GPT-4o (2024-08-06)1045
GPT-4o1036
GPT-4 0125 Preview1034
Mistral Large 2 (Instruct 2407)1029
Llama 3.1 405B Instruct1022
Gemini 1.5 Pro Experimental 08271007
Gemini 1.5 Pro Preview 0514994
GPT-4 Turbo (1106 Preview)992
DeepSeek-V3985
Llama 3.2 90B Vision Instruct984
Opus 3959
Gemini 1.5 Flash Preview 0514943
Gemini 1.5 Pro Preview 0409891
Claude 3 Sonnet879
Llama 3 70B Instruct871
Mistral Large 1.0811
Gemini 1.0 Pro685
Code Llama Instruct 34B598
Loading Atlas data…