Atlas

Benchmarks

← All benchmarks

BigCodeBench-Hard Instruct

Code · 2024-06-18

BigCodeBench evaluates code generation on practical programming tasks that compose calls across many libraries, scored by test execution as Pass@1. The Instruct split of the curated 148-task Hard subset, confirmed by the reported score lattice. Complete and Instruct are two prompt protocols over the same task pool, and the Hard subset is nested inside the full set, so all four scores share one factor group: they measure one construct and would otherwise supply four correlated votes for it.

Top models (higher is better)

ModelScore
Claude 3.7 Sonnet31.8
GPT-4.1 (2025-04-14)31.8
GPT-4.1 Mini (2025-04-14)31.8
DeepSeek-R129.7
Gemini 2.5 Pro Experimental 03-2529.7
GPT-4 Turbo29.1
Qwen2.5-Max29.1
DeepSeek-V328.4
GPT-4.1 Nano (2025-04-14)28.4
Llama 3.3 70B Instruct28.4
DeepSeek-V3-032427.7
Gemini Exp-120627.7
GPT-4o (2024-11-20)27.7
Llama 4 Maverick27.7
Qwen2.5 Coder 32B Instruct27.7
DeepSeek-V2.5-121027.0
Gemini 1.5 Pro Experimental 082727.0
Haiku 3.525.7
Claude 3.5 Sonnet (June 2024)25.7
Claude 3.5 Sonnet (Oct 2024)25.7
Gemini 1.5 Pro Experimental 080125.0
GPT-4o (2024-08-06)25.0
QwQ-32B-Preview25.0
DeepSeek-Coder-V2-Instruct24.3
Gemini 2.0 Flash Thinking Experimental 121924.3
Gemini Experimental 112124.3
Phi-424.3
DeepSeek-R1-Distill-Qwen-32B23.6
Gemini 2.0 Flash 00123.6
Grok Beta23.6
DeepSeek-V2.523.0
Gemini 2.0 Flash Experimental23.0
Llama 3.1 70B Instruct23.0
Opus 322.3
Llama 3.1 405B Instruct22.3
Llama 3 70B Instruct22.3
Mistral Large 2 (Instruct 2407)22.3
Qwen2.5 Instruct 32B22.3
DeepSeek-V2-Chat21.6
Mistral Small 3 24B Instruct 250121.6
Loading Atlas data…