Atlas

Benchmarks

← All benchmarks

BigCodeBench Complete

Code · 2024-06-18

BigCodeBench evaluates code generation on practical programming tasks that compose calls across many libraries, scored by test execution as Pass@1. The Complete split gives the model a docstring-style function signature to finish over the full 1,140-task set. Complete and Instruct are two prompt protocols over the same task pool, and the Hard subset is nested inside the full set, so all four scores share one factor group: they measure one construct and would otherwise supply four correlated votes for it.

Top models (higher is better)

ModelScore
Gemini Exp-120662.4
DeepSeek-V362.2
Llama 4 Maverick61.4
Gemini 2.0 Flash Experimental59.9
DeepSeek-Coder-V2-Instruct59.7
Haiku 3.559.0
GPT-4o (2024-11-20)58.9
Claude 3.5 Sonnet (June 2024)58.6
GPT-4 Turbo58.2
Gemini Experimental 112158.1
Qwen2.5 Coder 32B Instruct58.0
Claude 3.5 Sonnet (Oct 2024)57.5
Llama 3.3 70B Instruct57.5
Opus 357.4
GPT-4 061357.2
Hermes 2 Theta Llama 3 70B55.6
Phi-455.4
DeepSeek-R1-Distill-Qwen-32B54.9
Llama 3.1 70B Instruct54.8
Llama 3 70B Instruct54.5
QwQ-32B-Preview54.4
Claude 3 Sonnet53.8
DeepSeek-V2.5-121053.2
Codestral 240552.5
Qwen2.5 Instruct 32B52.3
Qwen2.5 14B Instruct52.2
GPT-3.5 Turbo 012550.6
Mistral Small 3 24B Instruct 250150.4
Mixtral 8x22B Instruct50.2
Claude 3 Haiku50.1
DeepSeek-R1-Distill-Llama-70B49.9
DeepSeek-V2-Chat49.0
Qwen2.5-Coder-7B-Instruct48.8
Phi-3 Medium 128K Instruct48.7
DeepSeek R1 Distill Qwen 14B48.4
DeepSeek-Coder-V2-Lite-Instruct47.6
Athene-70B47.5
Yi-Large47.2
DeepSeek-Coder-Base 33B46.6
Mistral Small Instruct 240946.6
Loading Atlas data…