Atlas

Benchmarks

← All benchmarks

BigCodeBench-Hard Complete

Code · 2024-06-18

BigCodeBench evaluates code generation on practical programming tasks that compose calls across many libraries, scored by test execution as Pass@1. BigCodeBench-Hard is the curated 148-task subset that the maintainers select as most demanding; this is its Complete split. The 148-task denominator is confirmed by the reported score lattice. Complete and Instruct are two prompt protocols over the same task pool, and the Hard subset is nested inside the full set, so all four scores share one factor group: they measure one construct and would otherwise supply four correlated votes for it.

Top models (higher is better)

ModelScore
DeepSeek-R140.5
DeepSeek-V340.5
Gemini Exp-120640.5
Gemini 2.5 Pro Experimental 03-2536.5
GPT-4o (2024-08-06)36.5
DeepSeek-V3-032435.8
Qwen2.5-Max35.8
Claude 3.5 Sonnet (Oct 2024)35.1
GPT-4 Turbo35.1
Haiku 3.534.5
GPT-4o (2024-11-20)34.5
Claude 3.7 Sonnet33.8
Gemini 2.0 Flash 00133.8
Gemini 2.0 Flash Experimental33.8
GPT-4.1 (2025-04-14)33.8
Qwen2.5 Coder 32B Instruct33.8
Claude 3.5 Sonnet (June 2024)33.1
Gemini 1.5 Pro 00232.4
Gemini 1.5 Pro Experimental 082731.8
GPT-4.1 Mini (2025-04-14)31.8
GPT-4.1 Nano (2025-04-14)31.8
Gemini 2.0 Flash Thinking Experimental 121930.4
Gemini Experimental 112130.4
Llama 3.1 405B Instruct30.4
Phi-430.4
Opus 329.7
DeepSeek-Coder-V2-Instruct29.7
Gemini 1.5 Pro Experimental 080129.7
Grok Beta29.7
Mistral Large 2 (Instruct 2407)29.7
DeepSeek-R1-Distill-Qwen-32B29.1
DeepSeek-V2.529.1
Llama 4 Maverick29.1
Llama 3.3 70B Instruct28.4
QwQ-32B-Preview28.4
Llama 3.1 70B Instruct27.7
Llama 3 70B Instruct27.0
Qwen2.5 Instruct 32B27.0
Claude 3 Sonnet26.4
DeepSeek-V2.5-121025.7
Loading Atlas data…