BigCodeBench-Hard Instruct
Code · 2024-06-18
BigCodeBench evaluates code generation on practical programming tasks that compose calls across many libraries, scored by test execution as Pass@1. The Instruct split of the curated 148-task Hard subset, confirmed by the reported score lattice. Complete and Instruct are two prompt protocols over the same task pool, and the Hard subset is nested inside the full set, so all four scores share one factor group: they measure one construct and would otherwise supply four correlated votes for it.