BigCodeBench Instruct
Code · 2024-06-18
BigCodeBench evaluates code generation on practical programming tasks that compose calls across many libraries, scored by test execution as Pass@1. The Instruct split states each of the 1,140 tasks as a natural-language instruction instead of a signature to complete, which is a harder prompt for the same work. Complete and Instruct are two prompt protocols over the same task pool, and the Hard subset is nested inside the full set, so all four scores share one factor group: they measure one construct and would otherwise supply four correlated votes for it.