Atlas

Benchmarks

← All benchmarks

BigCodeBench Instruct

Code · 2024-06-18

BigCodeBench evaluates code generation on practical programming tasks that compose calls across many libraries, scored by test execution as Pass@1. The Instruct split states each of the 1,140 tasks as a natural-language instruction instead of a signature to complete, which is a harder prompt for the same work. Complete and Instruct are two prompt protocols over the same task pool, and the Hard subset is nested inside the full set, so all four scores share one factor group: they measure one construct and would otherwise supply four correlated votes for it.

Top models (higher is better)

ModelScore
DeepSeek-V350.0
Llama 4 Maverick49.7
Qwen2.5 Coder 32B Instruct49.0
GPT-4.1 Mini (2025-04-14)48.9
DeepSeek-V2.5-121048.6
DeepSeek-Coder-V2-Instruct48.2
GPT-4 Turbo48.2
GPT-4o (2024-11-20)48.0
Gemini Exp-120647.0
Llama 3.3 70B Instruct46.9
Claude 3.5 Sonnet (June 2024)46.8
Haiku 3.546.1
Llama 3.1 70B Instruct46.1
GPT-4 061346.0
Gemini 2.0 Flash Experimental45.9
Hermes 2 Theta Llama 3 70B45.6
Opus 345.5
Phi-445.5
Gemini Experimental 112145.4
Mistral Small 3 24B Instruct 250145.3
Qwen2.5 Instruct 32B45.0
Claude 3.5 Sonnet (Oct 2024)44.6
QwQ-32B-Preview44.6
DeepSeek-R1-Distill-Qwen-32B43.9
Llama 3 70B Instruct43.6
Claude 3 Sonnet42.7
Codestral 240541.8
Mixtral 8x22B Instruct40.6
DeepSeek-V2-Chat40.4
Qwen2.5-Coder-7B-Instruct40.4
Qwen2.5 14B Instruct39.8
Claude 3 Haiku39.4
GPT-3.5 Turbo 012539.1
DeepSeek R1 Distill Qwen 14B38.1
Yi-Large37.7
Phi-3 Medium 128K Instruct37.6
Qwen2.5 7B Instruct37.6
C4AI Command R 08-202437.1
Athene-70B36.8
DeepSeek-Coder-V2-Lite-Instruct36.8
Loading Atlas data…