Atlas

Benchmarks

← All benchmarks

HumanEval

Code · 2021-07-07

HumanEval is a code-generation benchmark of 164 hand-written Python programming problems, each with hidden unit tests. Models are scored by pass@k, the fraction of problems solved by generated code, most often reported as pass@1. Higher is better.

Top models (higher is better)

ModelScore
Claude 3.5 Sonnet (Oct 2024)93.7
Claude 3.5 Sonnet (June 2024)92.0
Haiku 3.588.1
Opus 384.9
Gemini 1.5 Pro84.1
Claude 3 Haiku75.9
Gemini 1.0 Ultra74.4
Gemini 1.5 Flash74.3
Claude 3 Sonnet73.0
Gemini 1.0 Pro67.7
GPT-467.0
Grok 163.2
Loading Atlas data…